[HN Gopher] S1: Simple Test-Time Scaling
___________________________________________________________________
S1: Simple Test-Time Scaling
Author : t55
Score : 22 points
Date : 2025-02-03 17:56 UTC (5 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| randomcatuser wrote:
| really cool! i wonder what happens when you teach it to use tools
| inside the reasoning as well. could be even better!
| mncharity wrote:
| > This is similar to the "Superficial Alignment Hypothesis"
| presented in LIMA (Zhou et al., 2023), where the authors find
| that 1,000 examples can be sufficient to align a model to adhere
| to user preferences.
|
| Link: _LIMA: Less Is More for Alignment_
| https://proceedings.neurips.cc/paper_files/paper/2023/file/a... ,
| 1k cites:
| https://scholar.google.com/scholar?cites=1642843440474691780...
| artifishy_intel wrote:
| The diff between budget forcing on and off is all within the
| (surprisingly large) confidence intervals of evaluation datasets.
| Why spend more compute for no significant gain? Seems to distract
| from the high-value minimal reasoning ft set
|
| Also - In the main/first figure, why are r1 and o1 (the best
| performing models in Table 1) omitted?
|
| If you collect 59K and then pick the best 1K, is it really fair
| to say your approach is simple? Sifting through 59K examples
| doesn't seem simple.
|
| Good stuff though, cool to see how minimal we can get to distill
| good models (esp. at the manageable 32 size).
___________________________________________________________________
(page generated 2025-02-03 23:01 UTC)