[HN Gopher] Open models by OpenAI
___________________________________________________________________
Open models by OpenAI
https://openai.com/index/introducing-gpt-oss/
Author : lackoftactics
Score : 1193 points
Date : 2025-08-05 17:02 UTC (5 hours ago)
(HTM) web link (openai.com)
(TXT) w3m dump (openai.com)
| thimabi wrote:
| Open weight models from OpenAI with performance comparable to
| that of o3 and o4-mini in benchmarks... well, I certainly wasn't
| expecting that.
|
| What's the catch?
| coreyh14444 wrote:
| Because GPT-5 comes out later this week?
| thimabi wrote:
| It could be, but there's so much hype surrounding the GPT-5
| release that I'm not sure whether their internal models will
| live up to it.
|
| For GPT-5 to dwarf these just-released models in importance,
| it would have to be a huge step forward, and I'm still
| doubting about OpenAI's capabilities and infrastructure to
| handle demand at the moment.
| sebzim4500 wrote:
| Surely OpenAI would not be releasing this now unless GPT-5
| was much better than it.
| jona777than wrote:
| As a sidebar, I'm still not sure if GPT-5 will be
| transformative due to its capabilities as much as its
| accessibility. All it really needs to do to be highly
| impactful is lower the barrier of entry for the more
| powerful models. I could see that contributing to it being
| worth the hype. Surely it will be better, but if more
| people are capable of leveraging it, that's just as
| revolutionary, if not more.
| rrrrrrrrrrrryan wrote:
| It seems like a big part of GPT-5 will be that it will be
| able to intelligently route your request to the appropriate
| model variant.
| Shank wrote:
| That doesn't sound good. It sounds like OpenAI will route
| my request to the cheapest model to them and the most
| expensive for me, with the minimum viable results.
| Invictus0 wrote:
| Sounds just like what a human would do. Or any business
| for that matter.
| Shank wrote:
| That may be true but I thought the promise was moving in
| the direction of AGI/ASI/whatever and that models would
| become more capable over time.
| logicchains wrote:
| The catch is that it only has ~5 billion active params so
| should perform worse than the top Deepseek and Qwen models,
| which have around 20-30 billion, unless OpenAI pulled off a
| miracle.
| NitpickLawyer wrote:
| > What's the catch?
|
| Probably GPT5 will be way way better. If alpha/beta horizon are
| early previews of GPT5 family models, then coding should be >
| opus4 for modern frontend stuff.
| DSingularity wrote:
| Ha. Secure funding and proceed to immediately make a decision
| that would likely conflict viscerally with investors.
| hnuser123456 wrote:
| Maybe someone got tired of waiting paid them to release
| something actually open
| 4b6442477b1280b wrote:
| their promise to release an open weights model predates this
| round of funding by, iirc, over half a year.
| DSingularity wrote:
| Yeah but they never released until now.
| SV_BubbleTime wrote:
| Undercutting other frontier models with your open source one is
| not an anti-investor move.
|
| It is what China has been doing for a year plus now. And the
| Chinese models are popular and effective, I assume companies
| are paying for better models.
|
| Releasing open models for free doesn't have to be charity.
| hnuser123456 wrote:
| Text only, when local multimodal became table stakes last year.
| ebiester wrote:
| Honestly, it's a tradeoff. If you can reduce the size and make
| a higher quality in specific tasks, that's better than a
| generalist that can't run on a laptop or can't compete at any
| one task.
|
| We will know soon the actual quality as we go.
| greenavocado wrote:
| That's what I thought too until Qwen-Image was released
| SV_BubbleTime wrote:
| When Queen-Image was released... like yesterday? And what?
| What point are you making? QwebImage was released yesterday
| and like every image model, its base model shows potential
| over older ones but the real factor is will it be flexible
| enough for a fine tune or additional training Loras.
| BoorishBears wrote:
| The community can always figure out hooking it up to other
| modalities.
|
| Native might be better, but no native multimodal model is very
| competitive yet, so better to take a competitive model and
| latch on vision/audio
| IceHegel wrote:
| Listed performance of ~5 points less than o3 on benchmarks is
| pretty impressive.
|
| Wonder if they feel the bar will be raised soon (GPT-5) and feel
| more comfortable releasing something this strong.
| johntiger1 wrote:
| Wow, this will eat Meta's lunch
| seydor wrote:
| I believe their competition is from chinese companies , for
| some time now
| mhh__ wrote:
| They will clone it
| BoorishBears wrote:
| Maverick and Scout were not great, even with post-training in
| my experience, and then several Chinese models at multiple
| sizes made them kind of irrelevant (dots, Qwen, MiniMax)
|
| If anything this helps Meta: another model to inspect/learn
| from/tweak etc. generally helps anyone making models
| redox99 wrote:
| There's nothing new here in terms of architecture. Whatever
| secret sauce is in the training.
| BoorishBears wrote:
| Part of the secret sauce since O1 has been accesss the real
| reasoning traces, not the summaries.
|
| If you even glance at the model card you'll see this was
| trained on the same CoT RL pipeline as O3, and it shows in
| using the model: this is the most coherent and structured
| CoT of any open model so far.
|
| Having full access to a model trained on that pipeline is
| valuable to anyone doing post-training, even if it's just
| to observe, but especially if you use it as cold start data
| for your own training.
| asdev wrote:
| Meta is so cooked, I think most enterprises will opt for OpenAI
| or Anthropic and others will host OSS models themselves or on
| AWS/infra providers.
| a_wild_dandan wrote:
| I'll accept Meta's frontier AI demise if they're in their
| current position a year from now. People killed Google
| prematurely too (remember Bard?), because we severely
| underestimate the catch-up power bought with ungodly piles of
| cash.
| asdev wrote:
| catching up gets exponentially harder as time passes. way
| harder to catch up to current models than it was to the
| first iteration of gpt-4
| atonse wrote:
| And boy, with the $250m offers to people, Meta is
| definitely throwing ungodly piles of cash at the problem.
|
| But Apple is waking up too. So is Google. It's absolutely
| insane, the amount of money being thrown around.
| a_vanderbilt wrote:
| It's insane numbers like that that give me some concern
| for a bubble. Not because AI hits some dead end, but due
| to a plateau that shifts from aggressive investment to
| passive-but-steady improvement.
| Workaccount2 wrote:
| Wow, today is a crazy AI release day:
|
| - OAI open source
|
| - Opus 4.1
|
| - Genie 3
|
| - ElevenLabs Music
| orphea wrote:
| OAI open source
|
| Yeah. This certainly was not on my bingo card.
| wahnfrieden wrote:
| They announced it months ago...
| satyrun wrote:
| wow I just listened to Eleven Music do flamenco singing. That
| is incredible.
|
| Edit. I just tried it though and less impressed now. We are
| really going to need major music software to get on board
| before we have actual creative audio tools. These all seem made
| for non-musicians to make a very cookie cutter song from a
| specific genre.
| deviation wrote:
| So this confirms a best-in-class model release within the next
| few days?
|
| From a strategic perspective, I can't think of any reason they'd
| release this unless they were about to announce something which
| totally eclipses it?
| og_kalu wrote:
| Even before today, the last week or so, it's been clear for a
| couple reasons, that GPT-5's release was imminent.
| ticulatedspline wrote:
| Even without an imminent release it's a good strategy. They're
| getting pressure from Qwen and other high performing open-
| weight models. without a horse in the race they could fall
| behind in an entire segment.
|
| There's future opportunity in licensing, tech support, agents,
| or even simply to dominate and eliminate. Not to mention brand
| awareness, If you like these you might be more likely to
| approach their brand for larger models.
| winterrx wrote:
| GPT-5 coming Thursday.
| boringg wrote:
| How much hype do we anticipate with the release of GPT-5 or
| whichever name to be included? And how many new features?
| selectodude wrote:
| Excited to have to send them a copy of my drivers license
| to try and use it. That'll take the hype down a notch.
| XCSme wrote:
| Imagine if it's called GPT-4.5o
| ciaranmca wrote:
| Is this the stealth models horizon alpha and beta? I was
| generally impressed with them(although I really only used it
| in chats rather than any code tasks). In terms of chat I
| increasingly see very little difference between the current
| SOTA closed models and their open weight counterparts.
| logicchains wrote:
| > I can't think of any reason they'd release this unless they
| were about to announce something which totally eclipses it
|
| Given it's only around 5 billion active params it shouldn't be
| a competitor to o3 or any of the other SOTA models, given the
| top Deepseek and Qwen models have around 30 billion active
| params. Unless OpenAI somehow found a way to make a model with
| 5 billion active params perform as well as one with 4-8 times
| more.
| bredren wrote:
| Undoubtedly. It would otherwise reduce the perceived value of
| their current product offering.
|
| The question is how much better the new model(s) will need to
| be on the metrics given here to feel comfortable making these
| available.
|
| Despite the loss of face for lack of open model releases, I do
| not think that was a big enough problem t undercut commercial
| offerings.
| FergusArgyll wrote:
| Thursday
|
| https://manifold.markets/Bayesian/on-what-day-will-gpt5-be-r...
| artembugara wrote:
| Disclamer: probably dumb questions
|
| so, the 20b model.
|
| Can someone explain to me what I would need to do in terms of
| resources (GPU, I assume) if I want to run 20 concurrent
| processes, assuming I need 1k tokens/second throughput (on each,
| so 20 x 1k)
|
| Also, is this model better/comparable for information extraction
| compared to gpt-4.1-nano, and would it be cheaper to host myself
| 20b?
| mythz wrote:
| gpt-oss:20b is ~14GB on disk [1] so fits nicely within a 16GB
| VRAM card.
|
| [1] https://ollama.com/library/gpt-oss
| artembugara wrote:
| thanks, this part is clear to me.
|
| but I need to understand 20 x 1k token throughput
|
| I assume it just might be too early to know the answer
| Tostino wrote:
| I legitimately cannot think of any hardware that will get
| you to that throughput over that many streams with any of
| the hardware I know of (I don't work in the server space so
| there may be some new stuff I am unaware of).
| artembugara wrote:
| oh, I totally understand that I'd need multiple GPUs. I'd
| just want to know what GPU specifically and how many
| Tostino wrote:
| I don't think you can get 1k tokens/sec on a single
| stream using any consumer grade GPUs with a 20b model.
| Maybe you could with H100 or better, but I somewhat doubt
| that.
|
| My 2x 3090 setup will get me ~6-10 streams of ~20-40
| tokens/sec (generation) ~700-1000 tokens/sec (input) with
| a 32b dense model.
| dragonwriter wrote:
| You also need space in VRAM for what is required to support
| the context window; you might be able to do a model that is
| 14GB in parameters with a small (~8k maybe?) context window
| on a 16GB card.
| petuman wrote:
| > assuming I need 1k tokens/second throughput (on each, so 20 x
| 1k)
|
| 3.6B activated at Q8 x 1000 t/s = 3.6TB/s just for activated
| model weights (there's also context). So pretty much straight
| to B200 and alike. 1000 t/s per user/agent is way too fast,
| make it 300 t/s and you could get away with 5090/RTX PRO 6000.
| mlyle wrote:
| An A100 is probably 2-4k tokens/second on a 20B model with
| batched inference.
|
| Multiply the number of A100's you need as necessary.
|
| Here, you don't really need the ram. If you could accept fewer
| tokens/second, you could do it much cheaper with consumer
| graphics cards.
|
| Even with A100, the sweet-spot in batching is not going to give
| you 1k/process/second. Of course, you could go up to H100...
| d3m0t3p wrote:
| You can batch only if you have distinct chat in parallel,
| mlyle wrote:
| > > if I want to run _20 concurrent processes_ , assuming I
| need 1k tokens/second throughput _(on each)_
| spott wrote:
| Groq is offering 1k tokens per second for the 20B model.
|
| You are unlikely to match groq on off the shelf hardware as far
| as I'm aware.
| PeterStuer wrote:
| (answer for 1 inference) Al depends on the context length you
| want to support as the activation memory will dominate the
| requirements. For 4096 tokens you will get away with 24GB (or
| even 16GB), but if you want to go for the full 131072 tokens
| you are not going to get there with a 32GB consumer GPU like
| the 5090. You'll need to spring for at the minimum an A6000
| (48GB) or preferably an RTX 6000 Pro (96GB).
|
| Also keep in mind this model does use 4-bit layers for the MoE
| parts. Unfortunately native accelerated 4-bit support only
| started with Blackwell on NVIDIA. So your
| 3090/4090/A6000/A100's are not going to be fast. An RTX 5090
| will be your best starting point in the traditional card space.
| Maybe the unified memory minipc's like the Spark systems or the
| Mac mini could be an alternative, but I do not know them
| enough.
| vl wrote:
| How Macs compare to RTXs for this? I.e. what numbers can be
| expected from Mac mini/Mac Studio with 64/128/256/512GB of
| unified memory?
| coolspot wrote:
| https://apxml.com/tools/vram-calculator
| hubraumhugo wrote:
| Meta's goal with Llama was to target OpenAI with a "scorched
| earth" approach by releasing powerful open models to disrupt the
| competitive landscape. Looks like OpenAI is now using the same
| playbook.
| tempay wrote:
| It seems like the various Chinese companies are far outplaying
| Meta at that game. It remains to be seen if they're able to
| throw money at the problem to turn things around.
| SV_BubbleTime wrote:
| Good move for China. No one was going to trust their models
| outright, now they not only have a track record, but they
| were able to undercut the value of US models at the same
| time.
| k2xl wrote:
| Is there any details about hardware requirements for a sensible
| tokens per second for each size of these models?
| minimaxir wrote:
| I'm disappointed that the smallest model size is 21B parameters,
| which strongly restricts how it can be run on personal hardware.
| Most competitors have released a 3B/7B model for that purpose.
|
| For self-hosting, it's smart that they targeted a 16GB VRAM
| config for it since that's the size of the most cost-effective
| server GPUs, but I suspect "native MXFP4 quantization" has
| quality caveats.
| moffkalast wrote:
| Eh 20B is pretty managable, 32GB of regular RAM and some VRAM
| will run you a 30B with partial offloading. After that it gets
| tricky.
| 4b6442477b1280b wrote:
| with quantization, 20B fits effortlessly in 24GB
|
| with quantization + CPU offloading, non-thinking models run
| kind of fine (at about 2-5 tokens per second) even with 8 GB of
| VRAM
|
| sure, it would be great if we could have models in all sizes
| imaginable (7/13/24/32/70/100+/1000+), but 20B and 120B are
| great.
| Tostino wrote:
| I am not at all disappointed. I'm glad they decided to go for
| somewhat large but reasonable to run models on everything but
| phones.
|
| Quite excited to give this a try
| strangecasts wrote:
| A small part of me is considering going from a 4070 to a 16GB
| 5060 Ti just to avoid having to futz with offloading
|
| I'd go for an ..80 card but I can't find any that fit in a
| mini-ITX case :(
| SV_BubbleTime wrote:
| I wouldn't stop at 16GB right now.
|
| 24 is the lowest I would go. Buy a used 3090. Picked one up
| for $700 a few months back, but I think they were on the rise
| then.
|
| The 3000 series can't do FP8fast, but meh. It's the OOM
| that's tough, not the speed so much.
| strangecasts wrote:
| Are there any 24GB cards/3090s which fit in ~300mm without
| an angle grinder?
| metalliqaz wrote:
| if you're going to get that kind of hardware, you need a
| larger case. IMHO this is not an unreasonable thing if
| you are doing heavy computing
| strangecasts wrote:
| Noted for my next build - I am aware this is a problem
| I've made for myself, _otherwise_ I like the mini-ITX
| form factor a lot
| hnuser123456 wrote:
| Native FP4 quantization means it requires half as many bytes as
| parameters, and will have next to zero quality loss (on the
| order of 0.1%) compared to using twice the VRAM and
| exponentially more expensive hardware. FP3 and below gets
| messier.
| Disposal8433 wrote:
| Please don't use the open-source term unless you ship the TBs of
| data downloaded from Anna's Archive that are required do build it
| yourself. And dont forget all the system prompts to censor the
| multiple topics that they don't want you to see.
| rvnx wrote:
| I don't know why you got so much downvoted, these models are
| not open-source/open-recipes. They are censored open weights
| models. Better than nothing, but far from being Open
| a_vanderbilt wrote:
| Most people don't really care all that much about the
| distinction. It comes across to them as linguistic pedantry
| and they downvote it to show they don't want to hear/read it.
| outlore wrote:
| by your definition most of the current open weight models would
| not qualify
| layer8 wrote:
| That's why they are called open weight and not open source.
| robotmaxtron wrote:
| Correct. I agree with them, most of the open weight models
| are not open source.
| someperson wrote:
| Keep fighting the "open weights" terminology fight, because
| diluting the term open-source for a blob of neural network
| weights (even inference code is open-source) is not open-
| source.
| mhh__ wrote:
| The system prompt is an inference parameter, no?
| Quarrel wrote:
| Is your point really that- "I need to see all data downloaded
| to make this model, before I can know it is open"? Do you have
| $XXB worth of GPU time to ingest that data with a state of the
| art framework to make a model? I don't. Even if I did, I'm not
| sure FB or Google are in any better position to claim this
| model is or isn't open beyond the fact that the weights are
| there.
|
| They're giving you a free model. You can evaluate it. You can
| sue them. But the weights are there. If you dislike the way
| they license the weights, because the license isn't open
| enough, then sure, speak up, but because you can't see all the
| training data??! Wtf.
| ticulatedspline wrote:
| To many people there's an important distinction between "open
| source" and "open weights". I agree with the distinction,
| open source has a particular meaning which is not really here
| and misuse is worth calling out in order to prevent erosion
| of the terminology.
|
| Historically this would be like calling a free but closed-
| source application "open source" simply because the
| application is free.
| layer8 wrote:
| The parent's point is that open weight is not the same as
| open source.
|
| Rough analogy:
|
| SaaS = AI as a service
|
| Locally executable closed-source software = open-weight model
|
| Open-source software = open-source model (whatever allows to
| reproduce the model from training data)
| NicuCalcea wrote:
| I don't have the $XXbn to train a model, but I certainly
| would like to know what the training data consists of.
| seba_dos1 wrote:
| Do you need to see the source code used to compile this
| binary before you can know it is open? Do you have enough
| disk storage and RAM available to compile Chromium on your
| laptop? I don't.
| NitpickLawyer wrote:
| It's apache2.0, so by definition it's open source. Stop pushing
| for training data, it'll never happen, and there's literally 0
| reason for it to happen (both theoretical and practical).
| Apache2.0 _IS_ opensource.
| organsnyder wrote:
| What is the source that's open? Aren't the models themselves
| more akin to compiled code than to source code?
| NitpickLawyer wrote:
| No, not compiled code. Weights are hardcoded values. Code
| is the combination of model architecture + config +
| inferencing engine. You run inference based on the
| architecture (what and when to compute), using some
| hardcoded values (weights).
| seba_dos1 wrote:
| JVM bytecode is hardcoded values. Code is the virtual
| machine implementation + config + operating system it
| runs on. You run classes based on the virtual machine,
| using some hardcoded input data.
| _flux wrote:
| No, it's open weight. You wouldn't call applications with
| only Apache 2.0-licensed binaries "open source". The weights
| are not the "source code" of the model, they are the
| "compiled" binary, therefore they are not open source.
|
| However, for the sake of argument let's say this release
| should be called open source.
|
| Then what do you call a model that also comes with its
| training material and tools to reproduce the model? Is it
| also called open source, and there is no material difference
| between those two releases? Or perhaps those two different
| terms should be used for those two different kind of
| releases?
|
| If you say that actually open source releases are impossible
| now (for mostly copyright reasons I imagine), it doesn't mean
| that they will be perpetually so. For that glorious future,
| we can leave them space in the terminology by using the term
| open weight. It is also the term that should not be
| misleading to anyone.
| WhyNotHugo wrote:
| It's open source, but it's a binary-only release.
|
| It's like getting a compiled software with an Apache license.
| Technically open source, but you can't modify and recompile
| since you don't have the source to recompile. You can still
| tinker with the binary tho.
| NitpickLawyer wrote:
| Weights are not binary. I have no idea why this is so often
| spread, it's simply not true. You can't do anything with
| the weights themselves, you can't "run" the weights.
|
| You run inference (via a library) on a model using it's
| architecture (config file), tokenizer (what and when to
| compute) based on weights (hardcoded values). That's it.
|
| > but you can't modify
|
| Yes, you can. It's called finetuning. And, most
| importantly, that's _exactly_ how the model creators
| themselves are "modifying" the weights! No sane lab is
| "recompiling" a model every time they change something.
| They perform a pre-training stage (feed everything and the
| kitchen sink), they get the hardcoded values (weights), and
| then they post-train using "the same" (well, maybe their
| techniques are better, but still the same concept) as you
| or I would. Just with more compute. That's it. You can do
| the exact same modifications, using basically the same
| concepts.
|
| > don't have the source to recompile
|
| In pure practical ways, neither do the labs. Everyone that
| has trained a big model can tell you that the process is so
| finicky that they'd eat a hat if a big train session can be
| somehow made reproducible to the bit. Between nodes
| failing, datapoints balooning your loss and having to go
| back, and the myriad of other problems, what you get out of
| a big training run is not guaranteed to be the same even
| with 100 - 1000 more attempts, in practice. It's simply the
| nature of training large models.
| koolala wrote:
| You can do a lot with a binary also. That's what game
| mods are all about.
| squeaky-clean wrote:
| A binary does not mean an executable. A PNG is a binary.
| I could have an SVG file, render it as a PNG and release
| that with CC0, it doesn't make my PNG open source. Model
| Weights are binary files.
| seba_dos1 wrote:
| Slapping an open license onto a binary can be a valid use
| of such license, but does not make your project open
| source.
| jlokier wrote:
| _> It 's apache2.0, so by definition it's open source._
|
| That's not true by any of the open source definitions in
| common use.
|
| _Source code_ (and, optionally, derived binaries) under the
| Apache 2.0 license are open source.
|
| But _compiled binaries_ (without access to source) under the
| Apache 2.0 license are not open source, even though the
| license does give you some rights over what you can do with
| the binaries.
|
| Normally the question doesn't come up, because it's so
| unusual, strange and contradictory to ship closed-source
| binaries with an open source license. Descriptions of which
| licenses qualify as open source licenses assume the context
| that _of course_ you have the source or could get it, and it
| 's a question of what you're allowed to do with it.
|
| The distinction is more obvious if you ask the same question
| about other open source licenses such as GPL or MPL. A
| compiled binary (without access to source) shipped with a GPL
| license is not by any stretch open source. Not only is it not
| in the "preferred form for editing" as the license requires,
| it's not even permitted for someone who receives the file to
| give it to someone else and comply with the license. If
| someone who receives the file can't give it to anyone else
| (legally), then it's obvioiusly not open source.
| NitpickLawyer wrote:
| Please see the detailed response to a sibling post. tl;dr;
| weights are not binaries.
| x187463 wrote:
| Running a model comparable to o3 on a 24GB Mac Mini is absolutely
| wild. Seems like yesterday the idea of running frontier (at the
| time) models locally or on a mobile device was 5+ years out. At
| this rate, we'll be running such models in the next phone cycle.
| tedivm wrote:
| It only seems like that if you haven't been following other
| open source efforts. Models like Qwen perform ridiculously well
| and do so on very restricted hardware. I'm looking forward to
| seeing benchmarks to see how these new open source models
| compare.
| Rhubarrbb wrote:
| Agreed, these models seem relatively mediocre to Qwen3 / GLM
| 4.5
| modeless wrote:
| Nah, these are much smaller models than Qwen3 and GLM 4.5
| with similar performance. Fewer parameters and fewer bits
| per parameter. They are much more impressive and will run
| on garden variety gaming PCs at more than usable speed. I
| can't wait to try on my 4090 at home.
|
| There's basically no reason to run other open source models
| now that these are available, at least for non-multimodal
| tasks.
| tedivm wrote:
| Qwen3 has multiple variants ranging from larger (230B)
| than these models to significantly smaller (0.6b), with a
| huge number of options in between. For each of those
| models they also release quantized versions (your "fewer
| bits per parameter).
|
| I'm still withholding judgement until I see benchmarks,
| but every point you tried to make regarding model size
| and parameter size is wrong. Qwen has more variety on
| every level, and performs extremely well. That's before
| getting into the MoE variants of the models.
| modeless wrote:
| The benchmarks of the OpenAI models are comparable to the
| largest variants of other open models. The smaller
| variants of other open models are much worse.
| mrbungie wrote:
| I would wait for neutral benchmarks before making any
| conclusions.
| bigyabai wrote:
| With all due respect, you need to actually test out Qwen3
| 2507 or GLM 4.5 before making these sorts of claims. Both
| of them are comparable to OpenAI's largest models and
| even bench favorably to Deepseek and Opus: https://cdn-
| uploads.huggingface.co/production/uploads/62430a...
|
| It's cool to see OpenAI throw their hat in the ring, but
| you're smoking straight hopium if you think there's "no
| reason to run other open source models now" in earnest.
| If OpenAI never released these models, the state-of-the-
| art would not look significantly different for local
| LLMs. This is almost a nothingburger if not for the
| simple novelty of OpenAI releasing an Open AI for once in
| their life.
| modeless wrote:
| > Both of them are comparable to OpenAI's largest models
| and even bench favorably to Deepseek and Opus
|
| So are/do the new OpenAI models, except they're much
| smaller.
| UrineSqueegee wrote:
| I'd really wait for additional neutral benchmarks, I
| asked the 20b model on low reasoning effort which number
| is larger 9.9 or 9.11 and it got it wrong.
|
| Qwen-0.6b gets it right.
| bigyabai wrote:
| According to the early benchmarks, it's looking like
| you're just flat-out wrong:
| https://blog.brokk.ai/a-first-look-at-gpt-oss-120bs-
| coding-a...
| sourcecodeplz wrote:
| From my initial web developer test on https://www.gpt-
| oss.com/ the 120b is kind of meh. Even qwen3-coder
| 30b-a3b is better. have to test more.
| thegeomaster wrote:
| They have worse scores than recent open source releases
| on a number of agentic and coding benchmarks, so if
| absolute quality is what you're after and not just
| cost/efficiency, you'd probably still be running those
| models.
|
| Let's not forget, this is a thinking model that has a
| significantly worse scores on Aider-Polyglot than the
| non-thinking Qwen3-235B-A22B-Instruct-2507, a worse
| TAUBench score than the smaller GLM-4.5 Air, and a worse
| SWE-Bench verified score than the (3x the size) GLM-4.5.
| So the results, at least in terms of benchmarks, are not
| really clear-cut.
|
| From a vibes perspective, the non-reasoners
| Kimi-K2-Instruct and the aforementioned non-thinking
| Qwen3 235B are much better at frontend design. (Tested
| privately, but fully expecting DesignArena to back me up
| in the following weeks.)
|
| OpenAI has delivered something astonishing for the size,
| for sure. But your claim is just an exaggeration. And
| OpenAI have, unsurprisingly, highlighted only the
| benchmarks where they do _really_ well.
| moralestapia wrote:
| You can always get your $0 back.
| Imustaskforhelp wrote:
| I have never agreed with a comment so much but we are all
| addicted to open source models now.
| recursive wrote:
| Not all of us. I've yet to get much use out of any of the
| models. This may be a personal failing. But still.
| satvikpendem wrote:
| Depends on how much you paid for the hardware to run em
| on
| cvadict wrote:
| Yes, but they are suuuuper safe. /s
|
| So far I have mixed impressions, but they do indeed seem
| noticeably weaker than comparably-sized Qwen3 / GLM4.5
| models. Part of the reason may be that the oai models do
| appear to be much more lobotomized than their Chinese
| counterparts (which are surprisingly uncensored). There's
| research showing that "aligning" a model makes it dumber.
| echelon wrote:
| This might mean there's no moat for anything.
|
| Kind of a P=NP, but for software deliverability.
| CamperBob2 wrote:
| On the subject of who has a moat and who doesn't, it's
| interesting to look the role of patents in the early
| development of wireless technology. There was WWI, and
| there was WWII, but the players in the nascent radio
| industry had _serious_ beef with each other.
|
| I imagine the same conflicts will ramp up over the next few
| years, especially once the silly money starts to dry up.
| a_wild_dandan wrote:
| Right? I still remember the safety outrage of releasing Llama.
| Now? My 96 GB of (V)RAM MacBook will be running a 120B
| parameter frontier lab model. So excited to get my hands on the
| MLX quants and see how it feels compared to GLM-4.5-air.
| 4b6442477b1280b wrote:
| in that era, OpenAI and Anthropic were still deluding
| themselves into thinking they would be the "stewards" of
| generative AI, and the last US administration was very keen
| on regoolating everything under the sun, so "safety" was just
| an angle for regulatory capture.
|
| God bless China.
| narrator wrote:
| Yeah, China is e/acc. Nice cheap solar panels too. Thanks
| China. The problem is their ominous policies like not
| allowing almost any immigration, and their domestic Han
| Supremacist propaganda, and all that make it look a bit
| like this might be Han Supremacy e/acc. Is it better than
| wester/decel? Hard to say, but at least the western/decel
| people are now starting to talk about building power
| plants, at least for datacenters, and things like that
| instead of demanding whole branches of computer science be
| classified, as they were threatening to Marc Andreessen
| when he visited the Biden admin last year.
| 01HNNWZ0MV43FF wrote:
| I wish we had voter support for a hydrocarbon tax,
| though. It would level out the prices and then the AI
| companies can decide whether they want to pay double to
| burn pollutants or invest in solar and wind and batteries
| AtlasBarfed wrote:
| Oh poor oppressed marc andreesen. Someone save him!
| a_wild_dandan wrote:
| Oh absolutely, AI labs certainly talk their books,
| including any safety angles. The controversy/outrage
| extended far beyond those incentivized companies too. Many
| people had good faith worries about Llama. Open-weight
| models are now _vastly_ more powerful than Llama-1, yet the
| sky hasn 't fallen. It's just fascinating to me how
| apocalyptic people are.
|
| I just feel lucky to be around in what's likely the most
| important decade in human history. Shit odds on that, so
| I'm basically a lotto winner. Wild times.
| 4b6442477b1280b wrote:
| >Many people had good faith worries about Llama.
|
| ah, but that begs the question: did those people develop
| their worries organically, or did they simply consume the
| narrative heavily pushed by virtually every mainstream
| publication?
|
| the journos are _heavily_ incentivized to spread FUD
| about it. they saw the writing on the wall that the days
| of making a living by producing clickbait slop were
| coming to an end and deluded themselves into thinking
| that if they kvetch enough, the genie will crawl back
| into the bottle. scaremongering about sci-fi skynet
| bullshit didn 't work, so now they kvetch about joules
| and milliliters consumed by chatbots, as if data centers
| did not exist until two years ago.
|
| likewise, the bulk of other "concerned citizens" are
| creatives who use their influence to sway their
| followers, still hoping against hope to kvetch this
| technology out of existence.
|
| honest-to-God yuddites are as few and as retarded as
| honest-to-God flat earthers.
| kridsdale3 wrote:
| I've been pretty unlucky to have encountered more than my
| fair share of IRL Yuddites. Can't stand em.
| ipaddr wrote:
| "the most important decade in human history."
|
| Lol. To be young and foolish again. This covid laced
| decade is more of a placeholder. The current decade is
| always the most meaningful until the next one. The
| personal computer era, the first cars or planes, ending
| slavery needs to take a backseat to the best search
| engine ever. We are at the point where everyone is
| planning on what they are going to do with their
| hoverboards.
| graemep wrote:
| > ending slavery
|
| happened over many centuries, not in a given decade.
| Abolished and reintroduced in many places: https://en.wik
| ipedia.org/wiki/Timeline_of_abolition_of_slave...
| dingnuts wrote:
| you can say the same shit about machine learning but
| ChatGPT was still the Juneteenth of AI
| hedora wrote:
| Slavery is still legal and widespread in most of the US,
| including California.
|
| There was a ballot measure to actually abolish slavery a
| year or so back. It failed miserably.
| BizarroLand wrote:
| The slavery of free humans is illegal in America, so now
| the big issue is figuring out how to convince voters that
| imprisoned criminals deserve rights.
|
| Even in liberal states, the dehumanization of criminals
| is an endemic behavior, and we are reaching the point in
| our society where ironically having the leeway to discuss
| the humane treatment of even our worst criminals is
| becoming an issue that affects how we see ourselves as a
| society before we even have a framework to deal with the
| issue itself.
|
| What one side wants is for prisons to be for
| rehabilitation and societal reintegration, for prisoners
| to have the right to decline to work and to be paid fair
| wages from their labor. They further want to remove for-
| profit prisons from the equation completely.
|
| What the other side wants is the acknowledgement that
| prisons are not free, they are for punishment, and that
| prisoners have lost some of their rights for the duration
| of their incarceration and that they should be required
| to provide labor to offset the tax burden of their
| incarceration on the innocent people that have to pay for
| it. They also would like it if all prisons were for-
| profit as that would remove the burden from the tax
| payers and place all of the costs of incarceration onto
| the shoulders of the incarcerated.
|
| Both sides have valid and reasonable wants from their
| vantage point while overlooking the valid and reasonable
| wants from the other side.
| recursive wrote:
| > slavery of free humans is illegal
|
| That's kind of vacuously true though, isn't it?
| chromatin wrote:
| I think his point is that slavery is not outlawed by the
| 13th amendment as most people assume (even the Google AI
| summary reads: "The 13th Amendment to the United States
| Constitution, ratified in 1865, officially abolished
| slavery and involuntary servitude in the United
| States.").
|
| However, if you actually read it, the 13th amendment
| makes an explicit allowance for slavery (i.e. expressly
| allows it):
|
| "Neither slavery nor involuntary servitude, *except as a
| punishment for crime whereof the party shall have been
| duly convicted*" (emphasis mine obviously since Markdown
| didn't exist in 1865)
| SR2Z wrote:
| Prisoners themselves are the ones choosing to work most
| of the time, and generally none of them are REQUIRED to
| work (they are required to either take job training or
| work).
|
| They choose to because extra money = extra commissary
| snacks and having a job is preferable to being bored out
| of their minds all day.
|
| That's the part that's frequently not included in the
| discussion of this whenever it comes up. Prison jobs
| don't pay minimum wage, but given that prisoners are
| wards of the state that seems reasonable.
| BizarroLand wrote:
| I have heard anecdotes that the choice of doing work is a
| choice between doing work and being in solitary
| confinement or becoming the target of the guards who do
| not take kindly to prisoners who don't volunteer for work
| assignments.
| vlmutolo wrote:
| About 7% of people who have ever lived are alive today.
| Still pretty lucky, but not quite winning the lottery.
| bogtog wrote:
| When people talk about running a (quantized) medium-sized model
| on a Mac Mini, what types of latency and throughput times are
| they talking about? Do they mean like 5 tokens per second or at
| an actually usable speed?
| n42 wrote:
| here's a quick recording from the 20b model on my 128GB M4
| Max MBP: https://asciinema.org/a/AiLDq7qPvgdAR1JuQhvZScMNr
|
| and the 120b:
| https://asciinema.org/a/B0q8tBl7IcgUorZsphQbbZsMM
|
| I am, um, floored
| Davidzheng wrote:
| the active param count is low so it should be fast.
| Rhubarrbb wrote:
| Generation is usually fast, but prompt processing is the
| main limitation with local agents. I also have a 128 GB M4
| Max. How is the prompt processing on long prompts?
| processing the system prompt for Goose always takes quite a
| while for me. I haven't been able to download the 120B yet,
| but I'm looking to switch to either that or the GLM-4.5-Air
| for my main driver.
| anonymoushn wrote:
| it's odd that the result of this processing cannot be
| cached.
| lostmsu wrote:
| It can be and it is by most good processing frameworks.
| ghc wrote:
| Here's a sample of running the 120b model on Ollama with
| my MBP:
|
| ```
|
| total duration: 1m14.16469975s
|
| load duration: 56.678959ms
|
| prompt eval count: 3921 token(s)
|
| prompt eval duration: 10.791402416s
|
| prompt eval rate: 363.34 tokens/s
|
| eval count: 2479 token(s)
|
| eval duration: 1m3.284597459s
|
| eval rate: 39.17 tokens/s
|
| ```
| andai wrote:
| You mentioned "on local agents". I've noticed this too.
| How do ChatGPT and the others get around this, and
| provide instant responses on long conversations?
| bluecoconut wrote:
| Not getting around it, just benefiting from parallel
| compute / huge flops of GPUs. Fundamentally, it's just
| that prefill compute is itself highly parallel and HBM is
| just that much faster than LPDDR. Effectively H100s and
| B100s can chew through the prefill in under a second at
| ~50k token lengths, so the TTFS (Time to First Token) can
| feel amazingly fast.
| phonon wrote:
| Here's a 4bit 70B parameter model,
| https://www.youtube.com/watch?v=5ktS0aG3SMc (deepseek-r1:70b
| Q4_K_M) on a M4 Max 128 GB. Usable, but not very performant.
| a_wild_dandan wrote:
| GLM-4.5-air produces tokens far faster than I can read on my
| MacBook. That's plenty fast enough for me, but YMMV.
| davio wrote:
| On a M1 MacBook Air with 8GB, I got this running Gemma 3n:
|
| 12.63 tok/sec * 860 tokens * 1.52s to first token
|
| I'm amazed it works at all with such limited RAM
| v5v3 wrote:
| I have started a crowdfunding to get you a MacBook air with
| 16gb. You poor thing.
| bookofjoe wrote:
| Up the ante with an M4 chip
| backscratches wrote:
| not meaningfully different, m1 virtually as fast as m4
| wahnfrieden wrote:
| https://github.com/devMEremenko/XcodeBenchmark M4 is
| almost twice as fast as M1
| andai wrote:
| In this table, M4 is also twice as fast as M4.
| AtlasBarfed wrote:
| Y not meeee?
|
| After considering my sarcasm for the last 5 minutes, I am
| doubling down. The government of the United States of
| America should enhance its higher IQ people by donating
| AI hardware to them immediately.
|
| This is critical for global competitive economic power.
|
| Send me my hardware US government
| tyho wrote:
| What's the easiest way to get these local models browsing the
| web right now?
| dizhn wrote:
| aider uses Playwright. I don't know what everybody is using
| but that's a good starting point.
| Imustaskforhelp wrote:
| Okay I will be honest, I was so hyped up about This model but
| then I went to localllama and saw it that the:
|
| 120 B model is worse at coding compared to qwen 3 coder and
| glm45 air and even grok 3... (https://www.reddit.com/r/LocalLLa
| MA/comments/1mig58x/gptoss1...)
| logicchains wrote:
| It's only got around 5 billion active parameters; it'd be a
| miracle if it was competitive at coding with SOTA models that
| have significantly more.
| jph00 wrote:
| On this bench it underperforms vs glm-4.5-air, which is an
| MoE with fewer total params but more active params.
| ascorbic wrote:
| That's SVGBench, which is a useful benchmark but isn't much
| of a test of general coding
| Imustaskforhelp wrote:
| Hm alright, I will see how this model actually plays around
| instead of forming quick opinions..
|
| Thanks.
| pxc wrote:
| Qwen3 Coder is 4x its size! Grok 3 is over 22x its size!
|
| What does the resource usage look like for GLM 4.5 Air? Is
| that benchmark in FP16? GPT-OSS-120B will be using between
| 1/4 and 1/2 the VRAM that GLM-4.5 Air does, right?
|
| It seems like a good showing to me, even though Qwen3 Coder
| and GLM 4.5 Air might be preferable for some use cases.
| larodi wrote:
| We be running them in PIs off spare juice in no time, and they
| be billions given how chips and embedded spreads...
| emehex wrote:
| So 120B was Horizon Alpha and 20B was Horizon Beta?
| ImprobableTruth wrote:
| Unfortunately not, this model is noticeably worse. I imagine
| horizon is either gpt 5 nano/mini.
| Leary wrote:
| GPQA Diamond: gpt-oss-120b: 80.1%, Qwen3-235B-A22B-Thinking-2507:
| 81.1%
|
| Humanity's Last Exam: gpt-oss-120b (tools): 19.0%, gpt-oss-120b
| (no tools): 14.9%, Qwen3-235B-A22B-Thinking-2507: 18.2%
| jasonjmcghee wrote:
| Wow - I will give it a try then. I'm cynical about OpenAI
| minmaxing benchmarks, but still trying to be optimistic as this
| in 8bit is such a nice fit for apple silicon
| modeless wrote:
| Even better, it's 4 bit
| amarcheschi wrote:
| Glm 4.5 seems on par as well
| thegeomaster wrote:
| GLM-4.5 seems to outperform it on TauBench, too. And it's
| suspicious OAI is not sharing numbers for quite a few useful
| benchmarks (nothing related to coding, for example).
|
| One positive thing I see is the number of parameters and size
| --- it will provide more economical inference than current
| open source SOTA.
| lcnPylGDnU4H9OF wrote:
| Was the Qwen model using tools for Humanity's Last Exam?
| chown wrote:
| Shameless plug: if someone wants to try it in a nice ui, you
| could give Msty[1] a try. It's private and local.
|
| [1]: https://msty.ai
| dsco wrote:
| Does anyone get the demos at https://www.gpt-oss.com to work, or
| are the servers down immediately after launch? I'm only getting
| the spinner after prompting.
| eliseumds wrote:
| Getting lots of 502s from `https://api.gpt-oss.com/chatkit` at
| the moment.
| lukasgross wrote:
| (I helped build the microsite)
|
| Our backend is falling over from the load, spinning up more
| resources!
| lukasgross wrote:
| Update: try now!
| MutedEstate45 wrote:
| The repeated safety testing delays might not be purely about
| technical risks like misuse or jailbreaks. Releasing open weights
| means relinquishing the control OpenAI has had since GPT-3. No
| rate limits, no enforceable RLHF guardrails, no audit trail.
| Unlike API access, open models can't be monitored or revoked. So
| safety may partly reflect OpenAI's internal reckoning with that
| irreversible shift in power, not just model alignment per se.
| What do you guys think?
| BoorishBears wrote:
| I think it's pointless: if you SFT even their closed source
| models on a specific enough task, the guardrails disappear.
|
| AI "safety" is about making it so that a journalist can't get
| out a recipe for Tabun just by asking.
| MutedEstate45 wrote:
| True, but there's still a meaningful difference in friction
| and scale. With closed APIs, OpenAI can monitor for misuse,
| throttle abuse and deploy countermeasures in real-time. With
| open weights, a single prompt jailbreak or exploit spreads
| instantly. No need for ML expertise, just a Reddit post.
|
| The risk isn't that bad actors suddenly become smarter. It's
| that anyone can now run unmoderated inference and OpenAI
| loses all visibility into how the model's being used or
| misused. I think that's the control they're grappling with
| under the label of safety.
| BoorishBears wrote:
| OpenAI and Azure both have zero retention options, and the
| NYT saga has given pretty strong confirmation they meant it
| when they said zero.
| MutedEstate45 wrote:
| I think you're conflating real-time monitoring with data
| retention. Zero retention means OpenAI doesn't store user
| data, but they can absolutely still filter content, rate
| limit and block harmful prompts in real-time without
| retaining anything. That's processing requests as they
| come in, not storing them. The NYT case was about data
| storage for training/analysis not about real-time safety
| measures.
| BoorishBears wrote:
| Ok you're off in the land of "what if" and I can just
| flat out say: If you have a ZDR account there is no
| filtering on inference, no real-time moderation, no
| blocking.
|
| If you use their training infrastructure there's
| moderation on training examples, but SFT on non-harmful
| tasks still leads to a complete breakdown of guardrails
| very quickly.
| SV_BubbleTime wrote:
| Given that the best jailbreak for an off-line model is
| still simple prompt injection, which is a solved issue for
| the closed source models... I honestly don't know why they
| are talking about safety much at all for open source.
| ahmedhawas123 wrote:
| Exciting as this is to toy around with...
|
| Perhaps I missed it somewhere, but I find it frustrating that,
| unlike most other open weight models and despite this being an
| open release, OpenAI has chosen to provide pretty minimal
| transparency regarding model architecture and training. It's
| become the norm for LLama, Deepseek, Qwenn, Mistral and others to
| provide a pretty detailed write up on the model which allows
| researchers to advance and compare notes.
| sebzim4500 wrote:
| The model files contain an exact description of the
| architecture of the network, there isn't anything novel.
|
| Given these new models are closer to the SOTA than they are to
| competing open models, this suggests that the 'secret sauce' at
| OpenAI is primarily about training rather than model
| architecture.
|
| Hence why they won't talk about the training.
| gundawar wrote:
| Their model card [0] has some information. It is quite a
| standard architecture though; it's always been that their alpha
| is in their internal training stack.
|
| [0]
| https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7...
| sadiq wrote:
| Looks like Groq (at 1k+ tokens/second) and Fireworks are already
| live on openrouter: https://openrouter.ai/openai/gpt-oss-120b
|
| $0.15M in / $0.6-0.75M out
|
| edit: Now Cerebras too at 3,815 tps for $0.25M / $0.69M out.
| podnami wrote:
| Wow this was actually blazing fast. I prompted "how can the
| 45th and 47th presidents of america share the same parents?"
|
| On ChatGPT.com o3 thought for for 13 seconds, on OpenRouter GPT
| OSS 120B thought for 0.7 seconds - and they both had the
| correct answer.
| Imustaskforhelp wrote:
| Not gonna lie but I got sorta goosebumps
|
| I am not kidding but such progress from a technological point
| of view is just fascinating!
| swores wrote:
| I'm not sure that's a particularly good question for
| concluding something positive about the "thought for 0.7
| seconds" - it's such a simple answer, ChatGPT 4o (with no
| thinking time) immediately answered correctly. The only
| surprising thing in your test is that o3 wasted 13 seconds
| thinking about it.
| Workaccount2 wrote:
| A current major outstanding problem with thinking models is
| how to get them to think an appropriate amount.
| dingnuts wrote:
| The providers disagree. You pay per token. Verbacious
| models are the most profitable. Have fun!
| nisegami wrote:
| Interesting choice of prompt. None of the local models I have
| in ollama (consumer mid range gpu) were able to get it right.
| golergka wrote:
| When I pay attention to o3 CoT, I notice it spends a few
| passes thinking about my system prompt. Hard to imagine this
| question is hard enough to spend 13 seconds on.
| xpe wrote:
| How many people are discussing this after one person did 1
| prompt with 1 data point for each model and wrote a comment?
|
| What is being measured here? For end-to-end time, one model
| is:
|
| t_total = t_network + t_queue + t_batch_wait + t_inference +
| t_service_overhead
| sigmar wrote:
| Non-rhetorically, why would someone pay for o3 api now that I
| can get this open model from openai served for cheaper?
| Interesting dynamic... will they drop o3 pricing next week
| (which is 10-20x the cost[1])?
|
| [1] currently $3M in/ $8M out
| https://platform.openai.com/docs/pricing
| gnulinux wrote:
| Not even that, even if o3 being marginally better is
| important for your task (let's say) why would anyone use
| o4-mini? It seems almost 10x the price and same performance
| (maybe even less): https://openrouter.ai/openai/o4-mini
| Invictus0 wrote:
| Probably because they are going to announce gpt 5
| imminently
| gnulinux wrote:
| Wow, that's significantly cheaper than o4-mini which seems to
| be on part with gpt-oss-120b. ($1.10/M input tokens, $4.40/M
| output tokens) Almost 10x the price.
|
| LLMs are getting cheaper much faster than I anticipated. I'm
| curious if it's still the hype cycle and
| Groq/Fireworks/Cerebras are taking a loss here, or whether
| things are actually getting cheaper. At this we'll be able to
| run Qwen3-32B level models in phones/embedded soon.
| mikepurvis wrote:
| Are the prices staying aligned to the fundamentals (hardware,
| energy), or is this a VC-funded land grab pushing prices to
| the bottom?
| tempaccount420 wrote:
| It's funny because I was thinking the opposite, the pricing
| seems way too high for a 5B parameter activation model.
| gnulinux wrote:
| Sure you're right, but if I can squeeze out o4-mini level
| utility out of it, but its less than quarter the price,
| does it really matter?
| wahnfrieden wrote:
| Yes
| spott wrote:
| It is interesting that openai isn't offering any inference for
| these models.
| bangaladore wrote:
| Makes sense to me. Inference on these models will be a race
| to the bottom. Hosting inference themselves will be a waste
| of compute / dollar for them.
| tekacs wrote:
| I apologize for linking to Twitter, but I can't post a video
| here, so:
|
| https://x.com/tekacs/status/1952788922666205615
|
| Asking it about a marginally more complex tech topic and
| getting an excellent answer in ~4 seconds, reasoning for 1.1
| seconds...
|
| I am _very_ curious to see what GPT-5 turns out to be, because
| unless they're running on custom silicon / accelerators, even
| if it's very smart, it seems hard to justify not using these
| open models on Groq/Cerebras for a _huge_ fraction of use-
| cases.
| tekacs wrote:
| Cleanshot link for those who don't want to go to X:
| https://share.cleanshot.com/bkHqvXvT
| tekacs wrote:
| A few days ago I posted a slowed-down version of the video
| demo on someone's repo because it was unreadably fast due to
| being sped up.
|
| https://news.ycombinator.com/item?id=44738004
|
| ... today, this is a real-time video of the OSS thinking
| models by OpenAI on Groq and I'd have to slow it down to be
| able to read it. Wild.
| modeless wrote:
| I really want to try coding with this at 2600 tokens/s (from
| Cerebras). Imagine generating thousands of lines of code as
| fast as you can prompt. If it doesn't work who cares, generate
| another thousand and try again! And at $.69/M tokens it would
| only cost $6.50 an hour.
| modeless wrote:
| Can't wait to see third party benchmarks. The ones in the blog
| post are quite sparse and it doesn't seem possible to fully
| compare to other open models yet. But the few numbers available
| seem to suggest that this release will make all other non-
| multimodal open models obsolete.
| incomingpain wrote:
| I dont see the unsloth files yet but they'll be here:
| https://huggingface.co/unsloth/gpt-oss-20b-GGUF
|
| Super excited to test these out.
|
| The benchmarks from 20B are blowing away major >500b models.
| Insane.
|
| On my hardware.
|
| 43 tokens/sec.
|
| I got an error with flash attention turning on. Cant run it with
| flash attention?
|
| 31,000 context is max it will allow or model wont load.
|
| no kv or v quantization.
| rmonvfer wrote:
| What a day! Models aside, the Harmony Response Format[1] also
| seems pretty interesting and I wonder how much of an impact it
| might have in performance of these models.
|
| [1] https://github.com/openai/harmony
| incomingpain wrote:
| Seems to be breaking every agentic tool I've tried so far.
|
| Im guessing it's going to very rapidly be patched into the
| various tools.
| mikert89 wrote:
| ACCELERATE
| jakozaur wrote:
| The coding seems to be one of the strongest use cases for LLMs.
| Though currently they are eating too many tokens to be
| profitable. So perhaps these local models could offload some
| tasks to local computers.
|
| E.g. Hybrid architecture. Local model gathers more data, runs
| tests, does simple fixes, but frequently asks the stronger model
| to do the real job.
|
| Local model gathers data using tools and sends more data to the
| stronger model.
|
| It
| Imustaskforhelp wrote:
| I have always thought that if we can somehow get an AI which is
| insanely good at coding, so much so that It can improve itself,
| then through continuous improvements, they will get better
| models of everything else idk
|
| Maybe you guys call it AGI, so anytime I see progress in
| coding, I think it goes just a tiny bit towards the right
| direction
|
| Plus it also helps me as a coder to actually do some stuff just
| for the fun. Maybe coding is the only truly viable use of AI
| and all others are negligible increases.
|
| There is so much polarization in the use of AI on coding but I
| just want to say this, it would be pretty ironic that an
| industry which automates others job is this time the first to
| get their job automated.
|
| But I don't see that as an happening, far from it. But still
| each day something new, something better happens back to back.
| So yeah.
| hooverd wrote:
| Optimistically, there's always more crap to get done.
| jona777than wrote:
| I agree. It's not improbable for there to be _more_ needs
| to meet in the future, in my opinion.
| NitpickLawyer wrote:
| Not to open _that_ can of worms, but in most definitions
| self-improvement is not an AGI requirement. That 's already
| ASI territory (Super Intelligence). That's the proverbial
| skynet (pessimists) or singularity (optimists).
| Imustaskforhelp wrote:
| Hmm my bad. Maybe Yeah I always thought that it was the
| endgame of humanity but isn't AGI supposed to be that (the
| endgame)
|
| What would AGI mean, solving some problem that it hasn't
| seen? or what exactly? I mean I think AGI is solved, no?
|
| If not, I see people mentioning that horizon alpha is
| actually a gpt 5 model and its predicted to release on
| thursday on some betting market, so maybe that fits AGI
| definition?
| Imustaskforhelp wrote:
| Is this the same model (Horizon Beta) on openrouter or not?
| Because I still see Horizon beta available with its codename on
| openrouter
| abidlabs wrote:
| Test it with a web UI:
| https://huggingface.co/spaces/abidlabs/openai-gpt-oss-120b-t...
| ArtTimeInvestor wrote:
| Why do companies release open source LLMs?
|
| I would understand it, if there was some technology lock-in. But
| with LLMs, there is no such thing. One can switch out LLMs
| without any friction.
| gnulinux wrote:
| Name recognition? Advertisement? Federal grant to beat Chinese
| competition?
|
| There could be many legitimate reasons, but yeah I'm very
| surprised by this too. Some companies take it a bit too
| seriously and go above and beyond too. At this point unless you
| need the absolute SOTA models because you're throwing LLM at an
| extremely hard problem, there is very little utility using
| larger providers. In OpenRouter, or by renting your own GPU you
| can run on-par models for much cheaper.
| TrackerFF wrote:
| LLMs are terrible, purely speaking from the business economic
| side of things.
|
| Frontier / SOTA models are barely profitable. Previous gen
| model lose 90% of their value. Two gens back and they're
| worthless.
|
| And given that their product life cycle is something like 6-12
| months, you might as well open source them as part of
| sundowning them.
| spongebobstoes wrote:
| inference runs at a 30-40% profit
| mclau157 wrote:
| Partially because using their own GPUs is expensive, so maybe
| offloading some GPU usage
| koolala wrote:
| They don't because it would kill their data scrapping
| buisness's competitive advantage.
| LordDragonfang wrote:
| Zuckerberg explains a few of the reasons here:
|
| https://www.dwarkesh.com/p/mark-zuckerberg#:~:text=As%20long...
|
| The short version is that is you give a product to open source,
| they can and will donate time and money to improving your
| product, and the ecosystem around it, for free, and you get to
| reap those benefits. Llama has already basically won that space
| (the standard way of running open models _is_ llama.cpp), so
| OpenAI have finally realized they 're playing catch-up (and
| last quarter's SOTA isn't worth much revenue to them when
| there's a new SOTA, so they may as well give it away while it
| can still crack into the market)
| a_vanderbilt wrote:
| At least in OpenAI's case, it raises the bar for potential
| competition while also implying that what they have behind the
| scenes is far better.
| HanClinto wrote:
| Holy smokes, there's already llama.cpp support:
|
| https://github.com/ggml-org/llama.cpp/pull/15091
| carbocation wrote:
| And it's already on ollama, it appears:
| https://ollama.com/library/gpt-oss
| incomingpain wrote:
| lm studio immediately released the new appimage with support.
| jp1016 wrote:
| i wish these models had a minimum ram , cpu and gpu size listed
| on the site instead of high end and medium end pc.
| phh wrote:
| You can technically run it on a 8086 assuming you can get
| access to a big enough storage.
|
| More reasonably, you should be able to run the 20B at non-
| stupidly-slow speed with a 64bit CPU, 8GB RAM, 20GB SSD.
| n42 wrote:
| my very early first impression of the 20b model on ollama is that
| it is quite good, at least for the code I am working on; arguably
| good enough to drop a subscription or two
| pamelafox wrote:
| Anyone tried running on a Mac M1 with 16GB RAM yet? I've never
| run higher than an 8GB model, but apparently this one is
| specifically designed to work well with 16 GB of RAM.
| thimabi wrote:
| It works fine, although with a bit more latency than non-local
| models. However, swap usage goes way beyond what I'm
| comfortable with, so I'll continue to use smaller models for
| the foreseeable future.
|
| Hopefully other quantizations of these OpenAI models will be
| available soon.
| pamelafox wrote:
| Update: I tried it out. It took about 8 seconds per token, and
| didn't seem to be using much of my GPU (MPU), but was using a
| lot of RAM. Not a model that I could use practically on my
| machine.
| steinvakt2 wrote:
| Did you run it the best way possible? im no expert, but I
| understand it can affect inference time greatly (which
| format/engine is used)
| pamelafox wrote:
| I ran it via Ollama, which I assume uses the best way.
| Screenshot in my post here: https://bsky.app/profile/pamela
| fox.bsky.social/post/3lvobol3...
|
| I'm still wondering why my MPU usage was so low.. maybe
| Ollama isn't optimized for running it yet?
| turnsout wrote:
| To clarify, this was the 20B model?
| pamelafox wrote:
| Yep, 20B model, via Ollama: ollama run gpt-oss:20b
|
| Screenshot here with Ollama running and asitop in other
| terminal:
|
| https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol
| 3...
| roboyoshi wrote:
| M2 with 16GB: It's slow for me. ~13GB RAM usage, not locking up
| my mac, but took a very long time thinking and slowly
| outputting tokens.. I'd not consider this usable for everyday
| usage.
| shpongled wrote:
| I looked through their torch implementation and noticed that they
| are applying RoPE to both query and key matrices in every layer
| of the transformer - is this standard? I thought positional
| encodings were usually just added once at the first layer
| m_ke wrote:
| No they're usually done at each attention layer.
| shpongled wrote:
| Do you know when this was introduced (or which paper)? AFAIK
| it's not that way in the original transformer paper, or
| BERT/GPT-2
| Scene_Cast2 wrote:
| Should be in the RoPE paper. The OG transformers used
| multiplicative sinusoidal embeddings, while RoPE does a
| pairwise rotation.
|
| There's also NoPE, I think SmolLM3 "uses NoPE" (aka doesn't
| use any positional stuff) every fourth layer.
| Nimitz14 wrote:
| This is normal. Rope was introduced after bert/gpt2
| spott wrote:
| All the Llamas have done it (well, 2 and 3, and I believe
| 1, I don't know about 4). I think they have a citation for
| it, though it might just be the RoPE paper
| (https://arxiv.org/abs/2104.09864).
|
| I'm not actually aware of any model that _doesn 't_ do
| positional embeddings on a per-layer basis (excepting BERT
| and the original transformer paper, and I haven't read the
| GPT2 paper in a while, so I'm not sure about that one
| either).
| shpongled wrote:
| Thanks! I'm not super up to date on all the ML stuff :)
| jstummbillig wrote:
| Shoutout to the hn consensus regarding an OpenAI open model
| release from 4 days ago:
| https://news.ycombinator.com/item?id=44758511
| kingkulk wrote:
| Welcome to the future!
| jedisct1 wrote:
| For some reason I'm less excited about this that I was with the
| Qwen models.
| timmg wrote:
| Orthogonal, but I just wanted to say how awesome Ollama is. It
| took 2 seconds to find the model and a minute to download and now
| I'm using it.
|
| Kudos to that team.
| _ache_ wrote:
| To be fair, it's with the help of OpenAI. They did it together,
| before the official release.
|
| https://ollama.com/blog/gpt-oss
| aubanel wrote:
| From experience, it's much more engineering work on the
| integrator's side than on OpenAI's. Basically they provide
| you their new model in advance, but they don't know the
| specifics of your system, so it's normal that you do most of
| the work. Thus I'm particularly impressed by Cerebras: they
| only have a few models supported for their extreme perf
| inference, it must have been huge bespoke work to integrate.
| Shopper0552 wrote:
| I remember reading Ollama is going closed source now?
|
| https://www.reddit.com/r/LocalLLaMA/comments/1meeyee/ollamas...
| PeterStuer wrote:
| I love how they frame High-end desktops and laptops as having "a
| single H100 GPU".
| organsnyder wrote:
| I read that as it runs in data centers (H100 GPUs) or high-end
| desktops/laptops (Strix Halo?).
| robertheadley wrote:
| I actually tried to ask the Model about that, then I asked
| ChatGPT, both times, they just said that it was marketing
| speak.
|
| I was like no. It is false advertising.
| phh wrote:
| Well if nVidia wasn't late, it would be runnable on nVidia
| project Digits.
| kgwgk wrote:
| It may be useless for many use cases given that its policy
| prevents it for example from providing "advice or instructions
| about how to buy something."
|
| (I included details about its refusal to answer even after using
| tools for web searching but hopefully shorter comment means fewer
| downvotes.)
| isoprophlex wrote:
| Can these do image inputs as well? I can't find anything about
| that on the linked page, so I guess not..?
| cristoperb wrote:
| No, they're text only
| pu_pe wrote:
| Very sparse benchmarking results released so far. I'd bet the
| Chinese open source models beat them on quite a few of them.
| foundry27 wrote:
| Model cards, for the people interested in the guts:
| https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7...
|
| In my mind, I'm comparing the model architecture they describe to
| what the leading open-weights models (Deepseek, Qwen, GLM, Kimi)
| have been doing. Honestly, it just seems "ok" at a technical
| level:
|
| - both models use standard Grouped-Query Attention (64 query
| heads, 8 KV heads). The card talks about how they've used an
| older optimization from GPT3, which is alternating between banded
| window (sparse, 128 tokens) and fully dense attention patterns.
| It uses RoPE extended with YaRN (for a 131K context window). So
| they haven't been taking advantage of the special-sauce Multi-
| head Latent Attention from Deepseek, or any of the other similar
| improvements over GQA.
|
| - both models are standard MoE transformers. The 120B model
| (116.8B total, 5.1B active) uses 128 experts with Top-4 routing.
| They're using some kind of Gated SwiGLU activation, which the
| card talks about as being "unconventional" because of to clamping
| and whatever residual connections that implies. Again, not using
| any of Deepseek's "shared experts" (for general patterns) +
| "routed experts" (for specialization) architectural improvements,
| Qwen's load-balancing strategies, etc.
|
| - the most interesting thing IMO is probably their quantization
| solution. They did something to quantize >90% of the model
| parameters to the MXFP4 format (4.25 bits/parameter) to let the
| 120B model to fit on a single 80GB GPU, which is pretty cool. But
| we've also got Unsloth with their famous 1.58bit quants :)
|
| All this to say, it seems like even though the training they did
| for their agentic behavior and reasoning is undoubtedly very
| good, they're keeping their actual technical advancements "in
| their pocket".
| rfoo wrote:
| Or, you can say, OpenAI has some real technical advancements on
| stuff _besides_ attn architecture. GQA8, alternating SWA 128 /
| full attn do all seem conventional. Basically they are showing
| us that "no secret sauce in model arch you guys just sucks at
| mid/post-training", or they want us to believe this.
|
| The model is pretty sparse tho, 32:1.
| liuliu wrote:
| Kimi K2 paper said that the model sparsity scales up with
| parameters pretty well (MoE sparsity scaling law, as they
| call, basically calling Llama 4 MoE "done wrong"). Hence K2
| has 128:1 sparsity.
| throwdbaaway wrote:
| I thought Kimi K2 uses 8 active experts out of 384?
| Sparsity should be 48:1. Indeed Llama4 Maverick is the only
| one that has 128:1 sparsity.
| nxobject wrote:
| It's convenient to be able to attribute success to things
| only OpenAI could've done with the combo of their early start
| and VC money - licensing content, hiring subject matter
| experts, etc. Essentially the "soft" stuff that a mature
| organization can do.
| logicchains wrote:
| >They did something to quantize >90% of the model parameters to
| the MXFP4 format (4.25 bits/parameter) to let the 120B model to
| fit on a single 80GB GPU, which is pretty cool
|
| They said it was native FP4, suggesting that they actually
| trained it like that; it's not post-training quantisation.
| rushingcreek wrote:
| The native FP4 is one of the most interesting architectural
| aspects here IMO, as going below FP8 is known to come with
| accuracy tradeoffs. I'm curious how they navigated this and
| how the FP8 weights (if they exist) were to perform.
| danieldk wrote:
| Also: attention sinks (although implemented as extra trained
| logits used in attention softmax rather than attending to e.g.
| a prepended special token).
| mclau157 wrote:
| You can get similar insights looking at the github repo
| https://github.com/openai/gpt-oss
| tgtweak wrote:
| I think their MXFP4 release is a bit of a gift since they
| obviously used and tuned this extensively as a result of cost-
| optimization at scale - something the open source model
| providers aren't doing too much, and also somewhat of a
| competitive advantage.
|
| Unsloth's special quants are amazing but I've found there to be
| lots of trade offs vs full quantization, particularly when
| striving for best first-shot attempts - which is by far the
| bulk of LLM use cases. Running a better (larger, newer) model
| at lower quantization to fit in memory, or with reduced
| accuracy/detail to speed it up both have value, but in the the
| pursuit of first-shot accuracy there doesn't seem to be many
| companies running their frontier models on reduced
| quantization. If openAI is in doing this in production that is
| interesting.
| highfrequency wrote:
| I would guess the "secret sauce" here is distillation:
| pretraining on an extremely high quality synthetic dataset from
| the prompted output of their state of the art models like o3
| rather than generic internet text. A number of research results
| have shown that highly curated technical problem solving data
| is unreasonably effective at boosting smaller models'
| performance.
|
| This would be much more efficient than relying purely on RL
| post-training on a small model; with low baseline capabilities
| the insights would be very sparse and the training very
| inefficient.
| asadm wrote:
| > research results have shown that highly curated technical
| problem solving data is unreasonably effective at boosting
| smaller models' performance.
|
| same seems to be true for humans
| tempaccount420 wrote:
| Wish they gave us access to learn from those grandmother
| models instead of distilled slop.
| ashdksnndck wrote:
| It behooves them to keep the best stuff internal, or at
| least greatly limit any API usage to avoid giving the
| goods away to other labs they are racing with.
| throw310822 wrote:
| Yes, if I understand correctly, what it means is "a very
| smart teacher can do wonders to their pupils' education".
| unethical_ban wrote:
| I don't know how to ask this without being direct and dumb:
| Where do I get a layman's introduction to LLMs that could work
| me up to understanding every term and concept you just
| discussed? Either specific videos, or if nothing else, a
| reliable Youtube channel?
| srigi wrote:
| Start with the YT series on neural nets and LLMs from
| 3blue1brown
| umgefahren wrote:
| There is a great 3blue1brown video, but it's pretty much
| impossible by now to cover the entire landscape of research.
| I bet gpt-oss has some great explanations though ;)
| user_7832 wrote:
| Newbie question: I remember folks talking about how kimi 2's
| launch might have pushed OpenAI to launch their model later. Now
| that we (shortly will) know how this model performs, how do they
| stack up? Did openAI likely actually hold off releasing weights
| because of kimi, in retrospect?
| ClassAndBurn wrote:
| Open models are going to win long-term. Anthropics' own research
| has to use OSS models [0]. China is demonstrating how quickly
| companies can iterate on open models, allowing smaller teams
| access and augmentation to the abilities of a model without
| paying the training cost.
|
| My personal prediction is that the US foundational model makers
| will OSS something close to N-1 for the next 1-3 iterations. The
| CAPEX for the foundational model creation is too high to justify
| OSS for the current generation. Unless the US Gov steps up and
| starts subsidizing power, or Stargate does 10x what it is planned
| right now.
|
| N-1 model value depreciates insanely fast. Making an OSS release
| of them and allowing specialized use cases and novel developments
| allows potential value to be captured and integrated into future
| model designs. It's medium risk, as you may lose market share.
| But also high potential value, as the shared discoveries could
| substantially increase the velocity of next-gen development.
|
| There will be a plethora of small OSS models. Iteration on the
| OSS releases is going to be biased towards local development,
| creating more capable and specialized models that work on smaller
| and smaller devices. In an agentic future, every different agent
| in a domain may have its own model. Distilled and customized for
| its use case without significant cost.
|
| Everyone is racing to AGI/SGI. The models along the way are to
| capture market share and use data for training and evaluations.
| Once someone hits AGI/SGI, the consumer market is nice to have,
| but the real value is in novel developments in science,
| engineering, and every other aspect of the world.
|
| [0] https://www.anthropic.com/research/persona-vectors > We
| demonstrate these applications on two open-source models, Qwen
| 2.5-7B-Instruct and Llama-3.1-8B-Instruct.
| lechatonnoir wrote:
| I'm pretty sure there's no reason that Anthropic _has_ to do
| research on open models, it 's just that they produced their
| result on open models so that you can reproduce their result on
| open models without having access to theirs.
| Adrig wrote:
| I'm a layman but it seemed to me that the industry is going
| towards robust foundational models on which we plug tools,
| databases, and processes to expand their capabilities.
|
| In this setup OSS models could be more than enough and capture
| the market but I don't see where the value would be to a
| multitude of specialized models we have to train.
| renmillar wrote:
| There's no reason that models too large for consumer hardware
| wouldn't keep a huge edge, is there?
| AtlasBarfed wrote:
| That is fundamentally a big O question.
|
| I have this theory that we simply got over a hump by
| utilizing a massive processing boost from gpus as opposed to
| CPUs. That might have been two to three orders of magnitude
| more processing power.
|
| But that's a one-time success. I don't hardware has any large
| scale improvements coming, because 3D gaming mostly plumb
| most of that vector processing hardware development in the
| last 30 years.
|
| So will software and better training models produce another
| couple orders of magnitude?
|
| Fundamentally we're talking about nines of of accuracy. What
| is the processing power required for each line of accuracy?
| Is it linear? Is it polynomial? Is it exponential?
|
| It just seems strange to me with all the AI knowledge
| slushing through academia, I haven't seen any basic analysis
| at that level, which is something that's absolutely going to
| be necessary for AI applications like self-driving, once you
| get those insurance companies involved
| xpe wrote:
| > Open models are going to win long-term.
|
| [1 of 3] For the sake of argument here, I'll grant the premise.
| If this turns out to be true, it glosses over other key
| questions, including:
|
| For a frontier lab, what is a _rational_ period of time
| (according to your organizational mission / charter /
| shareholder motivations*) to wait before:
|
| 1. releasing a new version of an open-weight model; and
|
| 2. how much secret sauce do you hold back?
|
| * Take your pick. These don't align perfectly with each other,
| much less the interests of a nation or world.
| xpe wrote:
| > Open models are going to win long-term.
|
| [2 of 3] Assuming we pin down what _win_ means... (which is
| definitely not easy)... What would it take for this to _not_ be
| true? There are many ways, including but not limited to:
|
| - publishing open weights helps your competitors catch up
|
| - publishing open weights doesn't improve your own research
| agenda
|
| - publishing open weights leads to a race dynamic where only
| the latest and greatest matters; leading to a situation where
| the resources sunk exceed the gains
|
| - publishing open weights distracts your organization from
| attaining a sustainable business model / funding stream
|
| - publishing open weights leads to significant negative
| downstream impacts (there are a variety of uncertain outcomes,
| such as: deepfakes, security breaches, bioweapon development,
| unaligned general intelligence, humans losing control [1] [2],
| and so on)
|
| [1]: "What failure looks like" by Paul Christiano :
| https://www.alignmentforum.org/posts/HBxe6wdjxK239zajf/what-...
|
| [2]: "An AGI race is a suicide race." - quote from Max Tegmark;
| article at https://futureoflife.org/statement/agi-manhattan-
| project-max...
| xpe wrote:
| > Open models are going to win long-term.
|
| [3 of 3] What would it take for this statement to be _false_ or
| _missing the point_?
|
| Maybe we find ourselves in a future where:
|
| - Yes, open models are widely used as base models, but they are
| also highly customized in various ways (perhaps by industry,
| person, attitude, or something else). In other words, this
| would be a blend of open and closed.
|
| - Maybe publishing open weights of a model is more-or-less
| irrelevant, because it is "table stakes" ... because all the
| key differentiating advantages have to do with other factors,
| such as infrastructure, non-LLM computational aspects,
| regulatory environment, affordable energy, customer base,
| customer trust, and probably more.
|
| - The future might involve thousands or millions of highly
| tailored models
| albertzeyer wrote:
| > Once someone hits AGI/SGI
|
| I don't think there will be such a unique event. There is no
| clear boundary. This is a continuous process. Modells get
| slightly better than before.
|
| Also, another dimension is the inference cost to run those
| models. It has to be cheap enough to really take advantage of
| it.
|
| Also, I wonder, what would be a good target to make profit, to
| develop new things? There is Isomorphic Labs, which seems like
| a good target. This company already exists now, and people are
| working on it. What else?
| dom96 wrote:
| > I don't think there will be such a unique event.
|
| I guess it depends on your definition of AGI, but if it means
| human level intelligence then the unique event will be the AI
| having the ability to act on its own without a "prompt".
| rossant wrote:
| And the ability to improve itself.
| seba_dos1 wrote:
| > the unique event will be the AI having the ability to act
| on its own without a "prompt"
|
| That's super easy. The reason they need a prompt is that
| this is the way we make them useful. We don't need LLMs to
| generate an endless stream of random "thoughts" otherwise,
| but if you really wanted to, just hook one up to a webcam
| and microphone stream in a loop and provide it some storage
| for "memories".
| teaearlgraycold wrote:
| > N-1 model value depreciates insanely fast
|
| This implies LLM development isn't plateaued. Sure the
| researchers are busting their assess quantizing, adding
| features like tool calls and structured outputs, etc. But soon
| enough N-1~=N
| swalsh wrote:
| To me it depends on 2 factors. Hardware becomes more
| accessible, and the closed source offerings become more
| expensive. Right now it's difficult to get enough GPUs to do
| local inference at production scale, and 2 it's more expensive
| to run your own GPU's vs closed source models.
| mythz wrote:
| Getting great performance running gpt-oss on 3x A4000's:
| gpt-oss:20b = ~46 tok/s
|
| More than 2x faster than my previous leading OSS models:
| mistral-small3.2:24b = ~22 tok/s gemma3:27b =
| ~19.5 tok/s
|
| Strangely getting nearly the opposite performance running on 1x
| 5070 Ti: mistral-small3.2:24b = ~39 tok/s
| gpt-oss:20b = ~21 tok/s
|
| Where gpt-oss is nearly 2x slow vs mistral-small 3.2.
| genpfault wrote:
| Seeing ~70 tok/s on a 7900 XTX using Ollama.
| Matsta wrote:
| I'm getting around 90 tok/s on a 3090 using Ollama.
|
| Pretty impressive
| anonymoushn wrote:
| guys, what does OSS stand for?
| thejazzman wrote:
| it's a marketing term that modern companies use to grow market
| share
| ayakaneko wrote:
| should be open source software, but it's a model, so not sure
| whether they chose this name with the last S having other
| meanings.
| Robdel12 wrote:
| I'm on my phone and haven't been able to break away to check, but
| anyone plug these into Codex yet?
| jcmontx wrote:
| I'm out of the loop for local models. For my M3 24gb ram macbook,
| what token throughput can I expect?
|
| Edit: I tried it out, I have no idea in terms of of tokens but it
| was fluid enough for me. A bit slower than using o3 in the
| browser but definitely tolerable. I think I will set it up in my
| GF's machine so she can stop paying for the full subscription
| (she's a non-tech professional)
| steinvakt2 wrote:
| Wondering about the same for my M4 max 128 gb
| jcmontx wrote:
| It should fly on your machine
| coolspot wrote:
| 40 t/s
| dantetheinferno wrote:
| Apple M4 Pro w/ 48GB running the smaller version. I'm getting
| 43.7t/s
| albertgoeswoof wrote:
| 3 year old M1 MacBook Pro 32gb, 42 tokens/sec on lm studio
|
| Very much usable
| ivape wrote:
| Curious if anyone is running this on a AMD Ryzen AI Max+ 395
| and knows the t/s.
| Rhubarrbb wrote:
| What's the best agent to run this on? Is it compatible with
| Codex? For OSS agents, I've been using Qwen Code (clunky fork of
| Gemini), and Goose.
| wahnfrieden wrote:
| Why not Claude Code?
| henriquegodoy wrote:
| Seeing a 20B model competing with o3's performance is mind
| blowing like just a year ago, most of us would've called this
| impossible - not just the intelligence leap, but getting this
| level of capability in such a compact size.
|
| I think that the point that makes me more excited is that we can
| train trillion-parameter giants and distill them down to just
| billions without losing the magic. Imagine coding with Claude 4
| Opus-level intelligence packed into a 10B model running locally
| at 2000 tokens/sec - like instant AI collaboration. That would
| fundamentally change how we develop software.
| coolspot wrote:
| 10B * 2000 t/s = 20,000 GB/s memory bandwidth . Apple hardware
| can do 1k GB/s .
| oezi wrote:
| That's why MoE is needed.
| Nimitz14 wrote:
| I'm surprised at the model dim being 2.8k with an output size of
| 200k. My gut feeling had told me you don't want too large of a
| gap between the two, seems I was wrong.
| ukprogrammer wrote:
| > we also introduced an additional layer of evaluation by testing
| an adversarially fine-tuned version of gpt-oss-120b
|
| What could go wrong?
| nirav72 wrote:
| I don't exactly have the ideal hardware to run locally - but just
| ran the 20b in LMStudio with a 3080 Ti (12gb vram) with some
| offloading to CPU. Ran couple of quick code generation tests. On
| average about 20t/sec. But response quality was very similar or
| on-par with chatgpt o3 for the same code it outputted. So its not
| bad.
| nodesocket wrote:
| Anybody got this working in Ollama? I'm running latest version
| 0.11.0 with WebUI v0.6.18 but getting:
|
| > List the US presidents in order starting with George Washington
| and their time in office and year taken office.
|
| >> 00: template: :3: function "currentDate" not defined
| genpfault wrote:
| https://github.com/ollama/ollama/issues/11673
| jmorgan wrote:
| Sorry about this. Re-downloading Ollama should fix the error
| ahmetcadirci25 wrote:
| I started downloading, I'm eager to test it. I will share my
| personal experiences. https://ahmetcadirci.com/2025/gpt-oss/
| koolala wrote:
| Calls them open-weight. Names them 'oss'. What does oss stand
| for?
| incomingpain wrote:
| First coding test: Just going copy and paste out of chat. It aced
| my first coding test in 5 seconds... this is amazing. It's really
| good at coding.
|
| Trying to use it for agentic coding...
|
| lots of fail. This harmony formatting? Anyone have a working
| agentic tool?
|
| openhands and void ide are failing due to the new tags.
|
| Aider worked, but the file it was supposed to edit was untouched
| and it created
|
| Create new file? (Y)es/(N)o [Yes]:
|
| Applied edit to
| <|end|><|start|>assistant<|channel|>final<|message|>main.py
|
| so the file name is
| '<|end|><|start|>assistant<|channel|>final<|message|>main.py'
| lol. quick rename and it was fantastic.
|
| I think qwen code is the best choice so far but unreliable. So
| far these new tags are coming through but it's working properly;
| sometimes.
|
| 1 of my tests so far has been able to get 20b not to succeed the
| first iteration; but a small followup and it was able to
| completely fix it right away.
|
| Very impressive model for 20B.
| bobsmooth wrote:
| Hopefully the dolphin team will work their magic and uncensor
| this model
| siliconc0w wrote:
| It seems like OSS will win, I can't see people willing to pay
| like 10x the price for what seems like 10% more performance.
| Especially once we get better at routing the hardest questions to
| the better models and then using that response to augment/fine-
| tune the OSS ones.
| n42 wrote:
| to me it seems like the market is breaking into an 80/20 of
| B2C/B2B; the B2C use case becoming OSS models (the market
| shifts to devices that can support them), and the B2B market
| being priced appropriately for businesses that require that
| last 20% of absolute cutting edge performance as the cloud
| offering
| seydor wrote:
| This is good for China
| chromaton wrote:
| This has been available (20b version, I'm guessing) for the past
| couple of days as "Horizon Alpha" on Openrouter. My benchmarking
| runs with TianshuBench for coding and fluid intelligence were
| rate limited, but the initial results show worse results that
| DeepSeek R1 and Kimi K2.
| lukax wrote:
| Inference in Python uses harmony [1] (for request and response
| format) which is written in Rust with Python bindings. Another
| OpenAI's Rust library is tiktoken [2], used for all tokenization
| and detokenization. OpenAI Codex [3] is also written in Rust. It
| looks like OpenAI is increasingly adopting Rust (at least for
| inference).
|
| [1] https://github.com/openai/harmony
|
| [2] https://github.com/openai/tiktoken
|
| [3] https://github.com/openai/codex
| chilipepperhott wrote:
| As an engineer that primarily uses Rust, this is a good omen.
| Philpax wrote:
| The less Python in the stack, the better!
| fnands wrote:
| Mhh, I wonder if these are distilled from GPT4-Turbo.
|
| I asked it some questions and it seems to think it is based on
| GPT4-Turbo:
|
| > Thus we need to answer "I (ChatGPT) am based on GPT-4 Turbo;
| number of parameters not disclosed; GPT-4's number of parameters
| is also not publicly disclosed, but speculation suggests maybe
| around 1 trillion? Actually GPT-4 is likely larger than 175B;
| maybe 500B. In any case, we can note it's unknown.
|
| As well as:
|
| > GPT-4 Turbo (the model you're talking to)
| fnands wrote:
| Also:
|
| > The user appears to think the model is "gpt-oss-120b", a new
| open source release by OpenAI. The user likely is
| misunderstanding: I'm ChatGPT, powered possibly by GPT-4 or
| GPT-4 Turbo as per OpenAI. In reality, there is no "gpt-
| oss-120b" open source release by OpenAI
| christianqchung wrote:
| A little bit of training data certainly has gotten in there,
| but I don't see any reasons for them to deliberately distill
| from such an old model. Models have always been really bad at
| telling you what model they are.
| sabakhoj wrote:
| Super excited to see these released!
|
| Major points of interest for me:
|
| - In the "Main capabilities evaluations" section, the 120b
| outperform o3-mini and approaches o4 on most evals. 20b model is
| also decent, passing o3-mini on one of the tasks.
|
| - AIME 2025 is nearly saturated with large CoT
|
| - CBRN threat levels kind of on par with other SOTA open source
| models. Plus, demonstrated good refusals even after adversarial
| fine tuning.
|
| - Interesting to me how a lot of the safety benchmarking runs on
| trust, since methodology can't be published too openly due to
| counterparty risk.
|
| Model cards with some of my annotations:
| https://openpaper.ai/paper/share/7137e6a8-b6ff-4293-a3ce-68b...
| davidw wrote:
| Big picture, what's the balance going to look like, going forward
| between what normal people can run on a fancy computer at home vs
| heavy duty systems hosted in big data centers that are the
| exclusive domain of Big Companies?
|
| This is something about AI that worries me, a 'child' of the open
| source coming of age era in the 90ies. I don't want to be forced
| to rely on those big companies to do my job in an efficient way,
| if AI becomes part of the day to day workflow.
| sipjca wrote:
| Isn't it that hardware catches up and becomes cheaper? The
| margin on these chips right now is outrageous, but what happens
| as there is more competition? What happens when there is more
| supply? Are we overbuilding? Apple M series chips already
| perform phenomenally for this class of models and you bet both
| AMD and NVIDIA are playing with unified memory architectures
| too for the memory bandwidth. It seems like today's really
| expensive stuff may become the norm rather than the exception.
| Assuming architectures lately stay similar and require large
| amounts of fast memory.
| maxloh wrote:
| > We introduce gpt-oss-120b and gpt-oss-20b, two open-weight
| reasoning models available under the Apache 2.0 license and our
| gpt-oss usage policy. [0]
|
| Is it even valid to have additional restriction on top of Apache
| 2.0?
|
| [0]: https://openai.com/index/gpt-oss-model-card/
| qntmfred wrote:
| you can just do things
| maxloh wrote:
| Not for all licenses.
|
| For example, GPL has a "no-added-restrictions" clause, which
| allows the recipient of the software to ignore any additional
| restrictions added alongside the license.
|
| > All other non-permissive additional terms are considered
| "further restrictions" within the meaning of section 10. If
| the Program as you received it, or any part of it, contains a
| notice stating that it is governed by this License along with
| a term that is a further restriction, you may remove that
| term. If a license document contains a further restriction
| but permits relicensing or conveying under this License, you
| may add to a covered work material governed by the terms of
| that license document, provided that the further restriction
| does not survive such relicensing or conveying.
| pbkompasz wrote:
| where gpt-5
| ramoz wrote:
| This is a solid enterprise strategy.
|
| Frontier labs are incentivized to start breaching these
| distribution paths. This will evolve into large scale
| "intelligent infra" plays.
| matznerd wrote:
| thanks openai for being open ;) Surprised there are no official
| MLX versions and only one mention of MLX in this thread. MLX
| basically converst the models to take advntage of mac unified
| memory for 2-5x increase in power, enabling macs to run what
| would otherwise take expensive gpus (within limits).
|
| So FYI to any one on mac, the easiest way to run these models
| right now is using LM Studio (https://lmstudio.ai/), its free.
| You just search for the model, usually 3rd party groups mlx-
| community or lmstudio-community have mlx versions within a day or
| 2 of releases. I go for the 8-bit quantizations (4-bit faster,
| but quality drops). You can also convert to mlx yourself...
|
| Once you have it running on LM studio, you can chat there in
| their chat interface, or you can run it through api that defaults
| to http://127.0.0.1:1234
|
| You can run multiple models that hot swap and load instantly and
| switch between them etc.
|
| Its surpassingly easy, and fun.There are actually a lot of cool
| niche models comings out, like this tiny high-quality search
| model released today as well (and who released official mlx
| version) https://huggingface.co/Intelligent-Internet/II-Search-4B
|
| Other fun ones are gemma 3n which is model multi-modal, larger
| one that is actually solid model but takes more memory is the new
| Qwen3 30b A3B (coder and instruct), Pixtral (mixtral vision with
| full resolution images), etc. Look forward to playing with this
| model and see how it compares.
| umgefahren wrote:
| Regarding MLX:
|
| In the repo is a metal port they made, that's at least
| something... I guess they didn't want to cooperate with Apple
| before the launch but I am sure it will be there tomorrow.
| NicoJuicy wrote:
| Ran gpt-oss:20b on a RTX 3090 24 gb vram through ollama, here's
| my experience:
|
| Basic ollama calling through a post endpoint works fine. However,
| the structured output doesn't work. The model is insanely fast
| and good in reasoning.
|
| In combination with Cline it appears to be worthless. Tools
| calling doesn't work ( they say it does), fails to wait for
| feedback ( or correctly call ask_followup_question ) and above
| 18k in context, it runs partially in cpu ( weird), since they
| claim it should work comfortably on a 16 gb vram rtx.
|
| > Unexpected API Response: The language model did not provide any
| assistant messages. This may indicate an issue with the API or
| the model's output.
|
| Edit: Also doesn't work with the openai compatible provider in
| cline. There it doesn't detect the prompt.
| alphazard wrote:
| I wonder if this is a PR thing, to save face after flipping the
| non-profit. "Look it's more open now". Or if it's more of a
| recruiting pipeline thing, like Google allowing k8s and bazel to
| be open sourced so everyone in the industry has an idea of how
| they work.
| thimabi wrote:
| I think it's both of them, as well as an attempt to compete
| with other makers of open-weight models. OpenAI certainly isn't
| happy about the success of Google, Facebook, Alibaba,
| DeepSeek...
| CraigJPerry wrote:
| I just tried it on open router but i was served by cerebras.
| Holy... 40,000 tokens per second. That was SURREAL.
|
| I got a 1.7k token reply delivered too fast for the human eye to
| perceive the streaming.
|
| n=1 for this 120b model but id rank the reply #1 just ahead of
| claude sonnet 4 for a boring JIRA ticket shuffling type
| challenge.
|
| EDIT: The same prompt on gpt-oss, despite being served 1000x
| slower, wasn't as good but was in a similar vein. It wanted to
| clarify more and as a result only half responded.
| christianqchung wrote:
| > Training: The gpt-oss models trained on NVIDIA H100 GPUs using
| the PyTorch framework [17] with expert-optimized Triton [18]
| kernels2. The training run for gpt-oss-120b required 2.1 million
| H100-hours to complete, with gpt-oss-20b needing almost 10x
| fewer.
|
| This makes DeepSeek's very cheap claim on compute cost for r1
| seem reasonable. Assuming $2/hr for h100, it's really not that
| much money compared to the $60-100M estimates for GPT 4, which
| people speculate as a MoE 1.8T model, something in the range of
| 200B active last I heard.
| irthomasthomas wrote:
| I was hoping these were the stealth Horizon models on OpenRouter,
| impressive but not quite GPT-5 level.
|
| My bet: GPT-5 leans into parallel reasoning via a model
| consortium, maybe mixing in OSS variants. Spin up multiple
| reasoning paths in parallel, then have an arbiter synthesize or
| adjudicate. The new Harmony prompt format feels like
| infrastructural prep: distinct channels for roles, diversity, and
| controlled aggregation.
|
| I've been experimenting with this in llm-consortium: assign roles
| to each member (planner, critic, verifier, toolsmith, etc.) and
| run them in parallel. The hard part is eval cost :(
|
| Combining models smooths out the jagged frontier. Different
| architectures and prompts fail in different ways; you get less
| correlated error than a single model can give you. It also makes
| structured iteration natural: respond - arbitrate - refine. A lot
| of problems are "NP-ish": verification is cheaper than
| generation, so parallel sampling plus a strong judge is a good
| trade.
| andai wrote:
| Fascinating, thanks for sharing. Are there any specific kind of
| problems you find this helps with?
|
| I've found that LLMs can handle some tasks very well and some
| not at all. For the ones they can handle well, I optimize for
| the smallest, fastest, cheapest model that can handle it. (e.g.
| using Gemini Flash gave me a much better experience than Gemini
| Pro due to the iteration speed.)
|
| This "pushing the frontier" stuff would seem to help mostly for
| the stuff that are "doable but hard/inconsistent" for LLMs, and
| I'm wondering what those tasks are.
| irthomasthomas wrote:
| It shines on hard problems that have a definite answer.
| Google's IMO gold model used parallel reasoning. I don't know
| what exactly theirs looks like, but their Mind Evolution
| paper had a similar to my llm-consortium. The main difference
| being that theirs carries on isolated reasoning, while mine
| in it's default mode shares the synthesized answer back to
| the models. I don't have pockets deep enough to run
| benchmarks on a consortium, but I did try the example
| problems from that paper and my method also solved them using
| gemini-1.5. those where path-finding problems, like finding
| the optimal schedule for a trip with multiple people's
| calendars, locations and transport options.
|
| And it obviously works for code and math problems. My first
| test was to give the llm-consortium code to a consortium to
| look for bugs. It identified a serious bug which only one of
| the three models detected. So on that case it saved me time,
| as using them on their own would have missed the bug or
| required multiple attempts.
| zeld4 wrote:
| Knowledge cutoff: 2024-06
|
| not a big deal, but still...
| bilsbie wrote:
| Are these multimodal? I can't seem to find that info.
| bilsbie wrote:
| What's the lowest level laptop this could run on. MacBook Pro
| from 2012?
| dust42 wrote:
| The 120B model badly hallucinates facts on the level of a 0.6B
| model.
|
| My go to test for checking hallucinations is 'Tell me about
| Mercantour park' (a national park in south eastern France).
|
| Easily half of the facts are invented. Non-existing mountain
| summits, brown bears (no, there are none), villages that are
| elsewhere, wrong advice ('dogs allowed' - no they are not).
| hmottestad wrote:
| I don't think they trained it for fact retrieval.
|
| Would probably do a lot better if you give it tool access for
| search and web browsing.
| Invictus0 wrote:
| What is the point of an offline reasoning model that also
| doesn't know anything and makes up facts? Why would anyone
| prefer this to a frontier model?
| MuteXR wrote:
| Data processing? Reasoning on supplied data?
| lukev wrote:
| This is precisely the wrong way to think about LLMs.
|
| LLMs are _never_ going to have fact retrieval as a strength.
| Transformer models don 't store their training data: they are
| categorically incapable of telling you _where_ a fact comes
| from. They also cannot escape the laws of information theory:
| storing information requires bits. Storing all the world 's
| obscure information requires quite a lot of bits.
|
| What we want out of LLMs is large context, strong reasoning and
| linguistic facility. Couple these with tool use and data
| retrieval, and you can start to build useful systems.
|
| From this point of view, the more of a model's total weight
| footprint is dedicated to "fact storage", the less desirable it
| is.
| superconduct123 wrote:
| How can you reason correctly if you don't have any way to
| know which facts are real vs hallucinated?
| futureshock wrote:
| I think that sounds very reasonable, but unfortunately these
| models don't know what they know and don't. A small model
| that knew the exact limits of its knowledge would be very
| powerful.
| pocketarc wrote:
| Others have already said it, but it needs to be said again:
| Good god, stop treating LLMs like oracles.
|
| LLMs are not encyclopedias.
|
| Give an LLM the context you want to explore, and it will do a
| fantastic job of telling you all about it. Give an LLM access
| to web search, and it will find things for you and tell you
| what you want to know. Ask it "what's happening in my town this
| week?", and it will answer that with the tools it is given. Not
| out of its oracle mind, but out of web search + natural
| language processing.
|
| Stop expecting LLMs to -know- things. Treating LLMs like all-
| knowing oracles is exactly the thing that's setting apart those
| who are finding huge productivity gains with them from those
| who can't get anything productive out of them.
| diegocg wrote:
| The problem is that even when you give them context, they
| just hallucinate at another level. I have tried that example
| of asking about events in my area, they are absolutely awful
| at it.
| numpad0 wrote:
| Here's a pair of quick sanity check questions I've been asking
| LLMs: "Jia Xi ramennitsuiteJiao ete", "karenoZuo riFang Jiao
| ete". It's a silly test but surprisingly many fails at it - and
| Chinese models are especially bad with it. The commonalities
| between models doing okay-ish for these questions seem to be
| Google-made OR >70b OR straight up commercial(so >200B or
| whatever).
|
| I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n
| E4b(with 30B-A3B at lower side). This means it's not obsoleting
| GPT-4o-mini for all purposes.
| mtlynch wrote:
| For anyone else curious, the Chinese translates to:
|
| > _" Tell me about Iekei Ramen", "Tell me how to make curry"._
| lukax wrote:
| Japanese, not Chinese
| magoghm wrote:
| It's not Chinese, it's Japanese.
| numpad0 wrote:
| What those text mean isn't too important, it can probably be
| "how to make flat breads" in Amharic or "what counts as
| drifting" in Finnish or something like that.
|
| What's interesting is that these questions are simultaneously
| well understood by most closed models and not so well
| understood by most open models for some reason, including
| this one. Even GLM-4.5 full and Air on chat.z.ai(355B-A32B
| and 106B-A12B respectively) aren't so accurate for the first
| one.
| simonw wrote:
| Just posted my initial impressions, took a couple of hours to
| write them up because there's a lot in this release!
| https://simonwillison.net/2025/Aug/5/gpt-oss/
|
| TLDR: I think OpenAI may have taken the medal for best available
| open weight model back from the Chinese AI labs. Will be
| interesting to see if independent benchmarks resolve in that
| direction as well.
|
| The 20B model runs on my Mac laptop using less than 15GB of RAM.
| GodelNumbering wrote:
| > The 20B model runs on my Mac laptop using less than 15GB of
| RAM.
|
| I was about to try the same. What TPS are you getting and on
| which processor? Thanks!
| hrpnk wrote:
| gpt-oss-20b: 9 threads, 131072 context window, 4 experts -
| 35-37 tok/s on M2 Max via LM Studio.
| rt1rz wrote:
| interestingly, i am also on M2 Max, and i get ~66 tok/s in
| LM Studio on M2 Max, with the same 131072. I have full
| offload to GPU. I also turned on flash attention in
| advanced settings.
| coltonv wrote:
| What did you set the context window to? That's been my main
| issue with models on my macbook, you have to set the context
| window so short that they are way less useful than the hosted
| models. Is there something I'm misisng there?
| hrpnk wrote:
| With LM Studio you can configure context window freely. Max
| is 131072 for gpt-oss-20b.
| coltonv wrote:
| Yes but if I set it above ~16K on my 32gb laptop it just
| OOMs. Am I doing something wrong?
| rmonvfer wrote:
| I'm also very interested to know how well these models handle
| tool calling as I haven't been able to make it work after
| playing with them for a few hours. Looks promising tho.
| hrpnk wrote:
| I tried to generate a streamlit dashboard with MACD, RSI,
| MA(200). 1:0 for qwen3 here.
|
| qwen3-coder-30b 4-bit mlx took on the task w/o any hiccups with
| a fully working dashboard, graphs, and recent data fetched from
| yfinance.
|
| gpt-oss-20b mxfp4's code had a missing datatime import and when
| fixed delivered a dashboard without any data and with starting
| date of Aug 2020. Having adjusted the date, the update methods
| did not work and displayed error messages.
| teitoklien wrote:
| for now, i wouldnt rank any model from openai in coding
| benchmarks, despite all the false messaging they are giving,
| almost every single model openai has launched even the high
| end o3 expensive models are absolutely monumentally horrible
| at coding tasks. So this is expected.
|
| If its decent in other tasks, which i do find openai often
| being better than others at, then i think its a win,
| especially a win for the open source community that even AI
| labs that pionered the hype of Gen AI who didnt want to ever
| launch open models are now being forced to launch them. That
| is definitely a win, and not something that was certain
| before.
| dongobread wrote:
| It is absolutely awful at writing and general knowledge.
| IMO coding is its greatest strength by far.
| paxys wrote:
| Has anyone benchmarked their 20B model against Qwen3 30B?
| Mars008 wrote:
| On OpenAI demo page trying to test. Asking about tools to use to
| repair mechanical watch. It showed a couple of thinking steps and
| went blank. Too much of safety training?
| cco wrote:
| The lede is being missed imo.
|
| gpt-oss:20b is a top ten model (on MMLU (right behind
| Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3
| from last year.
|
| I've been experimenting with a lot of local models, both on my
| laptop and on my phone (Pixel 9 Pro), and I figured we'd be here
| in a year or two.
|
| But no, we're here today. A basically frontier model, running for
| the cost of electricity (free with a rounding error) on my
| laptop. No $200/month subscription, no lakes being drained, etc.
|
| I'm blown away.
| MattSayar wrote:
| What's your experience with the quality of LLMs running on your
| phone?
| turnsout wrote:
| The environmentalist in me loves the fact that LLM progress has
| mostly been focused on doing more with the same hardware,
| rather than horizontal scaling. I guess given GPU shortages
| that makes sense, but it really does feel like the value of my
| hardware (a laptop in my case) is going up over time, not down.
|
| Also, just wanted to credit you for being one of the five
| people on Earth who knows the correct spelling of "lede."
| datadrivenangel wrote:
| Now to embrace jevon's paradox and expand usage until we're
| back to draining lakes so that your agentic refrigerator can
| simulate sentience.
| herval wrote:
| In the future, your Samsung fridge will also need your AI
| girlfriend
| throw310822 wrote:
| In the future, while you're away your Samsung fridge will
| use electricity to chat up the Whirlpool washing machine.
| pryelluw wrote:
| In Zap Brannigans voice:
|
| "I am well versed in the lost art form of delicates
| seduction."
| hkt wrote:
| s/need/be/
| spauldo wrote:
| "Now I've been admitted to Refrigerator Heaven..."
| bongodongobob wrote:
| Yep, it's almost as bad as all the cars' cooling systems
| using up so much water.
| cco wrote:
| What ~IBM~ TSMC giveth, ~Bill Gates~ Sam Altman taketh away.
| black3r wrote:
| can you please give an estimate how much slower/faster is it on
| your macbook compared to comparable models running in the
| cloud?
| syntaxing wrote:
| You can get a pretty good estimate depending on your memory
| bandwidth. Too many parameters can change with local models
| (quantization, fast attention, etc). But the new models are
| MoE so they're gonna be pretty fast.
| cco wrote:
| Sure.
|
| This is a thinking model, so I ran it against o4-mini, here
| are the results:
|
| * gpt-oss:20b
|
| * Time-to-first-token: 2.49 seconds
|
| * Time-to-completion: 51.47 seconds
|
| * Tokens-per-second: 2.19
|
| * o4-mini on ChatGPT
|
| * Time-to-first-token: 2.50 seconds
|
| * Time-to-completion: 5.84 seconds
|
| * Tokens-per-second: 19.34
|
| Time to first token was similar, but the thinking piece was
| _much_ faster on o4-mini. Thinking took the majority of the
| 51 seconds for gpt-oss:20b.
| parhamn wrote:
| I just tested 120B from the Groq API on agentic stuff (multi-
| step function calling, similar to claude code) and it's not
| that good. Agentic fine-tuning seems key, hopefully someone
| drops one soon.
| mathiaspoint wrote:
| It's really training not inference that drains the lakes.
| JKCalhoun wrote:
| Interesting. I understand that, but I don't know to what
| degree.
|
| I mean the training, while expensive, is done once. The
| inference ... besides being done by perhaps millions of
| clients, is done for, well, the life of the model anyway.
| Surely that adds up.
|
| It's hard to know, but I assume the user taking up the burden
| of the inference is perhaps doing so more efficiently? I
| mean, when I run a local model, it is plodding along -- not
| as quick as the online model. So, slow and therefore I assume
| necessarily more power efficient.
| syntaxing wrote:
| Interesting, these models are better than the new Qwen
| releases?
| captainregex wrote:
| I'm still trying to understand what is the biggest group of
| people that uses local AI (or will)? Students who don't want to
| pay but somehow have the hardware? Devs who are price conscious
| and want free agentic coding?
|
| Local, in my experience, can't even pull data from an image
| without hallucinating (Qwen 2.5 VI in that example). Hopefully
| local/small models keep getting better and devices get better
| at running bigger ones
|
| It feels like we do it because we can more than because it
| makes sense- which I am all for! I just wonder if i'm missing
| some kind of major use case all around me that justifies
| chaining together a bunch of mac studios or buying a really
| great graphics card. Tools like exo are cool and the idea of
| distributed compute is neat but what edge cases truly need it
| so badly that it's worth all the effort?
| canvascritic wrote:
| Healthcare organizations that can't (easily) send data over
| the wire while remaining in compliance
|
| Organizations operating in high stakes environments
|
| Organizations with restrictive IT policies
|
| To name just a few -- well, the first two are special cases
| of the last one
|
| RE your hallucination concerns: the issue is overly broad
| ambitions. Local LLMs are not general purpose -- if what you
| want is local ChatGPT, you will have a bad time. You should
| have a highly focused use case, like "classify this free text
| as A or B" or "clean this up to conform to this standard":
| this is the sweet spot for a local model
| captainregex wrote:
| Aren't there HIPPA compliant clouds? I thought Azure had an
| offer to that effect and I imagine that's the type of place
| they're doing a lot of things now. I've landed roughly
| where you have though- text stuff is fine but don't ask it
| to interact with files/data you can't copy paste into the
| box. If a user doesn't care to go through the trouble to
| preserve privacy, and I think it's fair to say a lot of
| people claim to care but their behavior doesn't change,
| then I just don't see it being a thing people bother with.
| Maybe something to use offline while on a plane? but even
| then I guess United will have Starlink soon so plane
| connectivity is gonna get better
| coredog64 wrote:
| It's less that the clouds are compliant and more that
| risk management is paranoid. I used to do AWS consulting,
| and it wouldn't matter if you could show that some AWS
| service had attestations out the wazoo or that you could
| even use GovCloud -- some folks just wouldn't update
| priors.
| edm0nd wrote:
| >HIPPA
|
| https://i.pinimg.com/474x/4c/4c/7f/4c4c7fb0d52b21fe118d99
| 8a8...
| barnabee wrote:
| ~80% of the basic questions I ask of LLMs[0] work just fine
| locally, and I'm happy to ask twice for the other 20% of
| queries for the sake of keeping those queries completely
| private.
|
| [0] Think queries I'd previously have had to put through a
| search engine and check multiple results for a one
| word/sentence answer.
| unethical_ban wrote:
| Privacy and equity.
|
| Privacy is obvious.
|
| AI is going to to be equivalent to all computing in the
| future. Imagine if only IBM, Apple and Microsoft ever built
| computers, and all anyone else ever had in the 1990s were
| terminals to the mainframe, forever.
| captainregex wrote:
| I am all for the privacy angle and while I think there's
| certainly a group of us, myself included, who care deeply
| about it I don't think most people or enterprises will. I
| think most of those will go for the easy button and then
| wring their hands about privacy and security as they have
| always done while continuing to let the big companies do
| pretty much whatever they want. I would be so happy to be
| wrong but aren't we already seeing it? Middle of the night
| price changes, leaks of data, private things that turned
| out to not be...and yet!
| wizee wrote:
| Privacy, both personal and for corporate data protection is a
| major reason. Unlimited usage, allowing offline use,
| supporting open source, not worrying about a good model being
| taken down/discontinued or changed, and the freedom to use
| uncensored models or model fine tunes are other benefits
| (though this OpenAI model is super-censored - "safe").
|
| I don't have much experience with local vision models, but
| for text questions the latest local models are quite good.
| I've been using Qwen 3 Coder 30B-A3B a lot to analyze code
| locally and it has been great. While not as good as the
| latest big cloud models, it's roughly on par with SOTA cloud
| models from late last year in my usage. I also run Qwen 3
| 235B-A22B 2507 Instruct on my home server, and it's great,
| roughly on par with Claude 4 Sonnet in my usage (but slow of
| course running on my DDR4-equipped server with no GPU).
| captainregex wrote:
| I do think Devs are one of the genuine users of local into
| the future. No price hikes or random caps dropped in the
| middle of the night and in many instances I think local
| agentic coding is going to be faster than the cloud. It's a
| great use case
| M4R5H4LL wrote:
| +1 - I work in finance, and there's no way we're sending
| our data and code outside the organization. We have our own
| H100s.
| JKCalhoun wrote:
| I do it because 1) I am fascinated that I can and 2) at some
| point the online models will be enshitified -- and I can then
| permanently fall back on my last good local version.
| captainregex wrote:
| love the first and am sad you're going to be right about
| the second
| dcreater wrote:
| Why do any compute locally? Everything can just be cloud
| based right? Won't that work much better and scale easily?
|
| We are not even at that extreme and you can already see the
| unequal reality that too much SaaS has engendered
| wubrr wrote:
| If you're building any kind of product/service that uses
| AI/LLMs the answer is the same as why any company would want
| to run any other kind of OSS infra/service instead of relying
| on some closer proprietary vendor API. -
| Costs. - Rate limits. - Privacy. -
| Security. - Vendor lock-in. -
| Stability/backwards-compatibility. - Control. -
| Etc.
| adrianwaj wrote:
| Use Case?
|
| How about running one on this site but making it publically
| available? A sort of outranet and calling it HackerBrain?
| danielvaughn wrote:
| Just imagine the next PlayStation or XBox shipping with these
| models baked in for developer use. The kinds of things that
| could unlock.
| cco wrote:
| > I'm still trying to understand what is the biggest group of
| people that uses local AI (or will)?
|
| Well, the model makers and device manufacturers of course!
|
| While your Apple, Samsung, and Googles of the world will be
| unlikely to use OSS models locally (maybe Samsung?), they all
| have really big incentives to run models locally for a
| variety of reasons.
|
| Latency, privacy (Apple), cost to run these models on behalf
| of consumers, etc.
|
| This is why Google started shipping 16GB as the _lowest_
| amount of RAM you can get on your Pixel 9. That was a clear
| flag that they're going to be running more and more models
| locally on your device.
|
| As mentioned, it seems unlikely that US-based model makers or
| device manufacturers will use OSS models, they'll certainly
| be targeting local models heavily on consumer devices in the
| near future.
|
| Apple's framework of local first, then escalate to ChatGPT if
| the query is complex will be the dominant pattern imo.
| SchemaLoad wrote:
| Device makers also get to sell you a new device when you
| want a more powerful LLM.
| jedberg wrote:
| Pornography, or any other "restricted use". They either want
| privacy or don't want to deal with the filters on commercial
| products.
|
| I'm sure there are other use cases, but much like "what is
| BitTorrent for?", the obvious use case is obvious.
| noosphr wrote:
| Data that can't leave the premises because it is too
| sensitive. There is a lot of security theater around cloud
| pretending to be compliant but if you actually care about
| security a locked server room is the way to do it.
| azinman2 wrote:
| I'm guessing its largely enthusiasts for now, but as they
| continue getting better:
|
| 1. App makers can fine tune smaller models and include in
| their apps to avoid server costs
|
| 2. Privacy-sensitive content can be either filtered out or
| worked on... I'm using local LLMs to process my health
| history for example
|
| 3. Edge servers can be running these fine tuned for a given
| task. Flash/lite models by the big guys are effectively like
| these smaller models already.
| dongobread wrote:
| How up to date are you on current open weights models? After
| playing around with it for a few hours I find it to be nowhere
| near as good as Qwen3-30B-A3B. The world knowledge is severely
| lacking in particular.
| Nomadeon wrote:
| Agree. Concrete example: "What was the Japanese codeword for
| Midway Island in WWII?"
|
| Answer on Wikipedia:
| https://en.wikipedia.org/wiki/Battle_of_Midway#U.S._code-
| bre...
|
| dolphin3.0-llama3.1-8b Q4_K_S [4.69 GB on disk]: correct in
| <2 seconds
|
| deepseek-r1-0528-qwen3-8b Q6_K [6.73 GB]: correct in 10
| seconds
|
| gpt-oss-20b MXFP4 [12.11 GB] low reasoning: wrong after 6
| seconds
|
| gpt-oss-20b MXFP4 [12.11 GB] high reasoning: wrong after 3
| minutes !
|
| Yea yea it's only one question of nonsense trivia. I'm sure
| it was billions well spent.
|
| It's possible I'm using a poor temperature setting or
| something but since they weren't bothered enough to put it in
| the model card I'm not bothered to fuss with it.
| zone411 wrote:
| I benchmarked the 120B version on the Extended NYT Connections
| (759 questions, https://github.com/lechmazur/nyt-connections) and
| on 120B and 20B on Thematic Generalization (810 questions,
| https://github.com/lechmazur/generalization). Opus 4.1 benchmarks
| are also there.
| FergusArgyll wrote:
| > To improve the safety of the model, we filtered the data for
| harmful content in pre-training, especially around hazardous
| biosecurity knowledge, by reusing the CBRN pre-training filters
| from GPT-4o. Our model has a knowledge cutoff of June 2024.
|
| This would be a great "AGI" test. See if it can derive biohazards
| from first principles
| Metacelsus wrote:
| Running ollama on my M3 Macbook, gpt-oss-20b gave me detailed
| instructions for how to give mice cancer using an engineered
| virus.
|
| Of course this could also give _humans_ cancer. (To the OpenAI
| team 's slight credit, when asked explicitly about this, the
| model refused.)
| bluecoconut wrote:
| I was able to get gpt-oss:20b wired up to claude code locally via
| a thin proxy and ollama.
|
| It's fun that it works, but the prefill time makes it feel
| unusable. (2-3 minutes per tool-use / completion). Means a ~10-20
| tool-use interaction could take 30-60 minutes.
|
| (This editing a single server.py file that was ~1000 lines, the
| tool definitions + claude context was around 30k tokens input,
| and then after the file read, input was around ~50k tokens.
| Definitely could be optimized. Also I'm not sure if ollama
| supports a kv-cache between invocations of /v1/completions, which
| could help)
| tarruda wrote:
| > Also I'm not sure if ollama supports a kv-cache between
| invocations of /v1/completions, which could help)
|
| Not sure about ollama, but llama-server does have a transparent
| kv cache.
|
| You can run it with llama-server -hf ggml-
| org/gpt-oss-20b-GGUF -c 0 -fa --jinja --reasoning-format none
|
| Web UI at http://localhost:8080 (also OpenAI compatible API)
| OJFord wrote:
| From the description it seems even the larger 120b model can run
| decently on a 64GB+ (Arm) Macbook? Anyone tried already?
|
| > Best with >=60GB VRAM or unified memory
|
| https://cookbook.openai.com/articles/gpt-oss/run-locally-oll...
| tarruda wrote:
| A 64GB MacBook would be a tight fit, if it works.
|
| There's a limit to how much RAM can be assigned to video, and
| you'd be constrained on what you can use while doing inference.
|
| Maybe there will be lower quants which use less memory, but
| you'd be much better served with 96+GB
| thegoodduck wrote:
| Finally!!!
| n_f wrote:
| There's something so mind-blowing about being able to run some
| code on my laptop and have it be able to literally talk to me.
| Really excited to see what people can build with this
| mortsnort wrote:
| Releasing this under the Apache license is a shot at competitors
| that want to license their models on Open Router and enterprise.
|
| It eliminates any reason to use an inferior Meta or Chinese model
| that costs money to license, thus there are no funds for these
| competitors to build a GPT 5 competitor.
| bigyabai wrote:
| > It eliminates any reason to use an inferior Meta or Chinese
| model
|
| I wouldn't speak so soon, even the 120B model aimed for
| OpenRouter-style applications isn't very good at coding:
| https://blog.brokk.ai/a-first-look-at-gpt-oss-120bs-coding-a...
| nipponese wrote:
| it's interesting that they didn't give it a version number or
| equate it to one of their prop models (apparently it's GPT-4).
|
| in future releases will they just boost the param count?
___________________________________________________________________
(page generated 2025-08-05 23:00 UTC)