[HN Gopher] Tencent's 'Hunyuan-T1'-The First Mamba-Powered Ultra...
       ___________________________________________________________________
        
       Tencent's 'Hunyuan-T1'-The First Mamba-Powered Ultra-Large Model
        
       Author : marban
       Score  : 107 points
       Date   : 2025-03-22 17:25 UTC (5 hours ago)
        
 (HTM) web link (llm.hunyuan.tencent.com)
 (TXT) w3m dump (llm.hunyuan.tencent.com)
        
       | chis wrote:
       | Kobe?
        
       | nixpulvis wrote:
       | Some of the text is cut off while reading on my phone.
       | Embarrassing.
        
         | drysine wrote:
         | Don't be so harsh on your phone)
        
         | pkkkzip wrote:
         | thanks for sharing did you contact tencent support ?
        
         | jrflowers wrote:
         | Why are you embarrassed? You can always put your phone down and
         | read it on desktop later
        
       | notShabu wrote:
       | The romanization of these names is always confusing b/c stripped
       | of the character and tone it's just gibberish. "Hunyuan" or Hun
       | Yuan  in chinese means "Primordial Chaos" or "Original Unity".
       | 
       | This helps as more chinese products and services hit the market
       | and makes it easier to remember. The naming is similar to the
       | popularity of greek mythology in western products. (e.g. all the
       | products named "Apollo")
        
         | klabb3 wrote:
         | > The naming is similar to the popularity of greek mythology in
         | western products. (e.g. all the products named "Apollo")
         | 
         | Popular? So you're saying that all the VPs who have come up
         | with the mind bendingly unique and creative name Prometheus
         | didn't do so out of level 10 vision?
        
         | Y_Y wrote:
         | I think it's particularly egregious that they use such a lossy
         | encoding. I can't read the hanzi, but at least "Hun yuan" would
         | have been more helpful, or even "Hu4n yua1n" would have enabled
         | me to pronounce it or look it up without having the context to
         | guess which characters it was representing.
        
       | ttoinou wrote:
       | the excellent performance demonstrated by the models fully proves
       | the crucial role of reinforcement learning in the optimization
       | process
       | 
       | What if this reinforcement is just gaming the benchmarks
       | (Goodhart's law) without providing better answers elsewhere, how
       | would we notice it ?
        
         | m3kw9 wrote:
         | When actual people start using it
        
         | dartos wrote:
         | I mean all optimization algorithms do is game a benchmark.
         | That's the whole point.
         | 
         | The hard part is making the benchmark meaningful in the first
         | place.
        
           | TeMPOraL wrote:
           | Yeah, and if anything, RL has a rep of being _too good at
           | this job_ , because of all the cases where it gamed a
           | benchmark by picking up on some environmental factor the
           | supervisors hadn't thought of (numerical instabilities,
           | rounding, bugs, etc.).
        
           | einpoklum wrote:
           | No, that is patently false. Many optimization algorithms
           | which computer scientists, mathematicians or software
           | developers devise do not involve benchmakrs at all, and apply
           | to all possible inputs/instances of their respective
           | computational problems.
        
         | mentalgear wrote:
         | The trick is that the benchmarks must have a wide enough
         | distribution so that a well scoring model is potentially useful
         | for the widest span of users.
         | 
         | There also would need to be a guarantee (or checking of the
         | model somehow) that model providers don't just train on the
         | benchmarks. Solutions are dynamic components (random names,
         | numbers, etc) or private parts of benchmarks.
        
       | cowpig wrote:
       | Does the fact that they are linking to a Huggingface demo imply
       | they will be releasing the weights?
        
       | Magi604 wrote:
       | So many models coming out these days, so many developments
       | happening in the AI space in general, it's kinda hard to keep up
       | with it all. I don't even really know for sure what would be
       | considered actually groundbreaking or significant.
        
         | bicx wrote:
         | I try to generally keep up with the overall trends, but I'm an
         | engineer at a resource-constrained startup, not a research
         | scientist. I want to see real-world application, at least mid-
         | term value, minimum lock-in, and strong supportability. Until
         | then, I just don't have time to think about it.
        
           | squigz wrote:
           | You may both be interested in this newsletter
           | 
           | https://nlp.elvissaravia.com/t/ai
        
         | threeseed wrote:
         | For me nothing has been groundbreaking nor significant. What we
         | are seeing is the same in every new innovation, a suite of
         | micro-innovations which improves efficiency and reduces cost.
         | 
         | But LLMs are still fundamentally a stochastic parrot that
         | depends heavily on source data to produce useful results. So we
         | will go through a lull until there is some new groundbreaking
         | research which moves everything forward. And then the cycle
         | repeats.
        
       | kristianp wrote:
       | So their Large Model was 389b parameters, how big is their Ultra-
       | Large model?
        
       | sroussey wrote:
       | It's exciting to see a Mamba based model do so well.
        
       ___________________________________________________________________
       (page generated 2025-03-22 23:00 UTC)