[HN Gopher] Qwen2.5-Max: Exploring the intelligence of large-sca...
       ___________________________________________________________________
        
       Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
        
       Author : rochoa
       Score  : 92 points
       Date   : 2025-01-28 15:55 UTC (7 hours ago)
        
 (HTM) web link (qwenlm.github.io)
 (TXT) w3m dump (qwenlm.github.io)
        
       | ecshafer wrote:
       | A Chinese company announcing this on Spring Festival eve, that is
       | very surprising. The deep seek announcement must have put a fire
       | under them. I am surprised anything is being done right now in
       | these Chinese tech companies.
        
         | lostmsu wrote:
         | It's like when Gemini topped Chatbot Arena Leaderboard, and
         | OpenAI released a model next day.
        
           | lousken wrote:
           | is gemini really better than e.g. claude 3.5?
        
             | og_kalu wrote:
             | Mostly, but not for coding. Also the arena is pretty much a
             | vibes benchmark. Don't take it too seriously. Livebench is
             | a better indicator
        
             | jaggs wrote:
             | Gemini is actually pretty useless because of its rate
             | limits.
        
         | rfoo wrote:
         | Well, DeepSeek engineers are (desperately) fire-fighting as
         | they don't have nearly as much capacity as needed. Competitors
         | either already rushed release or decided to do an hush release
         | of whatever they had in the pipeline. Sounds like everyone is
         | working L
        
           | nostradumbasp wrote:
           | They are being attacked as well.
           | 
           | https://apnews.com/article/deepseek-ai-artificial-
           | intelligen...
        
       | bigcat12345678 wrote:
       | Party goes on
        
       | GaggiX wrote:
       | Now they need to finetune it like R1 and o1 and it will be
       | competitive with SOTA models.
        
       | simonw wrote:
       | This appears to be Qwen's new best model, API only for the
       | moment, which they say is better than DeepSeek v3.
        
         | rochoa wrote:
         | It is available at https://chat.qwenlm.ai/, under the model
         | selector.
        
       | BhavdeepSethi wrote:
       | HuggingFace demo: https://huggingface.co/spaces/Qwen/Qwen2.5-Max-
       | Demo
       | 
       | Source: https://x.com/Alibaba_Qwen/status/1884263157574820053
        
       | alecco wrote:
       | No weights, no proof.
        
         | Tiberium wrote:
         | Would you say the same for OpenAI releasing new models?
        
           | kragen wrote:
           | I may be misremembering, but I think he has.
        
       | jondwillis wrote:
       | The significance of _all_ of these releases at once is not lost
       | on me. But the reason for it is lost on me. Is there some
       | convention? Is this political? Business strategy?
        
         | logicchains wrote:
         | Today is the last day before the Chinese New Year.
        
           | voxgen wrote:
           | My thoughts go out to the poor engineers who got put on call
           | because someone scheduled a product release on the day before
           | the biggest holiday of their year.
        
       | a_wild_dandan wrote:
       | > We evaluate Qwen2.5-Max alongside leading models
       | 
       | > [...] we are unable to access the proprietary models such as
       | GPT-4o and Claude-3.5-Sonnet. Therefore, we evaluate Qwen2.5-Max
       | against DeepSeek V3
       | 
       | "We'll compare our proprietary model to other proprietary models.
       | Except when we don't. Then we'll compare to non-proprietary
       | models."
        
       | Jackson__ wrote:
       | >Many critical details regarding this scaling process were only
       | disclosed with the recent release of DeepSeek V3
       | 
       | And so they decide to not disclose their own training information
       | just after they told everyone how useful it was to get Deepseeks?
       | Honestly can't say I care about "nearly as good as o1" when its a
       | closed API with no additional info.
        
         | voxgen wrote:
         | It's not even "nearly as good as o1". They only compared to the
         | older 4o.
         | 
         | You can safely assume Qwen2.5-Max will score worse than all of
         | the recent reasoning models (o1, DeepSeek-R1, Gemini 2.0 Flash
         | Thinking).
         | 
         | It'll probably become a very strong model if/when they apply RL
         | training for reasoning. However, all the successful recipes for
         | this are closed source, so it may take some time. They could do
         | SFT based on another model's reasoning chains in the meantime,
         | though the DeepSeek-R1 technical report noted that it's not as
         | good as RL training.
        
       | kragen wrote:
       | I thought there were three DeepSeek items on the HN front page,
       | but this turned out to be a fourth one, because it's the Qwen
       | team saying they have a secret version of Qwen that's actually
       | better than DeepSeek-V3.
       | 
       | I don't remember the last time 20% of the HN front page was about
       | the same thing. Then again, _nobody_ remembers the last time a
       | company 's market cap fell by 569 billion dollars like NVIDIA did
       | yesterday.
        
         | caycep wrote:
         | it's a scaling law for stocks!
        
       | zone411 wrote:
       | I just ran my NYT Connections benchmark on it: 18.6, up from 14.8
       | for Qwen 2.5 72B. I'll run my other benchmarks later.
       | 
       | https://github.com/lechmazur/nyt-connections/
        
       | Havoc wrote:
       | Kinda ambivalent about MoE in cloud. Where it could really shine
       | though is in desktop class gear. Memory is starting to get fast
       | enough where we might see MoEs being not painfully slow soon for
       | large-ish models.
        
       ___________________________________________________________________
       (page generated 2025-01-28 23:01 UTC)