[HN Gopher] Tencent's 'Hunyuan-T1'-The First Mamba-Powered Ultra...
___________________________________________________________________
Tencent's 'Hunyuan-T1'-The First Mamba-Powered Ultra-Large Model
Author : marban
Score : 107 points
Date : 2025-03-22 17:25 UTC (5 hours ago)
(HTM) web link (llm.hunyuan.tencent.com)
(TXT) w3m dump (llm.hunyuan.tencent.com)
| chis wrote:
| Kobe?
| nixpulvis wrote:
| Some of the text is cut off while reading on my phone.
| Embarrassing.
| drysine wrote:
| Don't be so harsh on your phone)
| pkkkzip wrote:
| thanks for sharing did you contact tencent support ?
| jrflowers wrote:
| Why are you embarrassed? You can always put your phone down and
| read it on desktop later
| notShabu wrote:
| The romanization of these names is always confusing b/c stripped
| of the character and tone it's just gibberish. "Hunyuan" or Hun
| Yuan in chinese means "Primordial Chaos" or "Original Unity".
|
| This helps as more chinese products and services hit the market
| and makes it easier to remember. The naming is similar to the
| popularity of greek mythology in western products. (e.g. all the
| products named "Apollo")
| klabb3 wrote:
| > The naming is similar to the popularity of greek mythology in
| western products. (e.g. all the products named "Apollo")
|
| Popular? So you're saying that all the VPs who have come up
| with the mind bendingly unique and creative name Prometheus
| didn't do so out of level 10 vision?
| Y_Y wrote:
| I think it's particularly egregious that they use such a lossy
| encoding. I can't read the hanzi, but at least "Hun yuan" would
| have been more helpful, or even "Hu4n yua1n" would have enabled
| me to pronounce it or look it up without having the context to
| guess which characters it was representing.
| ttoinou wrote:
| the excellent performance demonstrated by the models fully proves
| the crucial role of reinforcement learning in the optimization
| process
|
| What if this reinforcement is just gaming the benchmarks
| (Goodhart's law) without providing better answers elsewhere, how
| would we notice it ?
| m3kw9 wrote:
| When actual people start using it
| dartos wrote:
| I mean all optimization algorithms do is game a benchmark.
| That's the whole point.
|
| The hard part is making the benchmark meaningful in the first
| place.
| TeMPOraL wrote:
| Yeah, and if anything, RL has a rep of being _too good at
| this job_ , because of all the cases where it gamed a
| benchmark by picking up on some environmental factor the
| supervisors hadn't thought of (numerical instabilities,
| rounding, bugs, etc.).
| einpoklum wrote:
| No, that is patently false. Many optimization algorithms
| which computer scientists, mathematicians or software
| developers devise do not involve benchmakrs at all, and apply
| to all possible inputs/instances of their respective
| computational problems.
| mentalgear wrote:
| The trick is that the benchmarks must have a wide enough
| distribution so that a well scoring model is potentially useful
| for the widest span of users.
|
| There also would need to be a guarantee (or checking of the
| model somehow) that model providers don't just train on the
| benchmarks. Solutions are dynamic components (random names,
| numbers, etc) or private parts of benchmarks.
| cowpig wrote:
| Does the fact that they are linking to a Huggingface demo imply
| they will be releasing the weights?
| Magi604 wrote:
| So many models coming out these days, so many developments
| happening in the AI space in general, it's kinda hard to keep up
| with it all. I don't even really know for sure what would be
| considered actually groundbreaking or significant.
| bicx wrote:
| I try to generally keep up with the overall trends, but I'm an
| engineer at a resource-constrained startup, not a research
| scientist. I want to see real-world application, at least mid-
| term value, minimum lock-in, and strong supportability. Until
| then, I just don't have time to think about it.
| squigz wrote:
| You may both be interested in this newsletter
|
| https://nlp.elvissaravia.com/t/ai
| threeseed wrote:
| For me nothing has been groundbreaking nor significant. What we
| are seeing is the same in every new innovation, a suite of
| micro-innovations which improves efficiency and reduces cost.
|
| But LLMs are still fundamentally a stochastic parrot that
| depends heavily on source data to produce useful results. So we
| will go through a lull until there is some new groundbreaking
| research which moves everything forward. And then the cycle
| repeats.
| kristianp wrote:
| So their Large Model was 389b parameters, how big is their Ultra-
| Large model?
| sroussey wrote:
| It's exciting to see a Mamba based model do so well.
___________________________________________________________________
(page generated 2025-03-22 23:00 UTC)