[HN Gopher] Reinforcement Learning from Human Feedback
       ___________________________________________________________________
        
       Reinforcement Learning from Human Feedback
        
       https://arxiv.org/abs/2504.12501
        
       Author : onurkanbkrc
       Score  : 97 points
       Date   : 2026-02-07 12:53 UTC (10 hours ago)
        
 (HTM) web link (rlhfbook.com)
 (TXT) w3m dump (rlhfbook.com)
        
       | klelatti wrote:
       | Web version with links, etc:
       | 
       | https://rlhfbook.com/
        
         | dang wrote:
         | Thanks! We've switched to that above from
         | https://arxiv.org/abs/2504.12501, and put the latter in the
         | toptext.
        
       | verdverm wrote:
       | Last time I saw Nathan say something about the book, he's
       | actively working on the next version and looking for feedback,
       | check his socials
        
         | leggerss wrote:
         | You could say he's also learning from human feedback
        
       | dang wrote:
       | Related. Others?
       | 
       |  _RLHF Book_ - https://news.ycombinator.com/item?id=42902936 -
       | Feb 2025 (37 comments)
        
       ___________________________________________________________________
       (page generated 2026-02-07 23:00 UTC)