[HN Gopher] Reinforcement Learning from Human Feedback
___________________________________________________________________
Reinforcement Learning from Human Feedback
https://arxiv.org/abs/2504.12501
Author : onurkanbkrc
Score : 97 points
Date : 2026-02-07 12:53 UTC (10 hours ago)
(HTM) web link (rlhfbook.com)
(TXT) w3m dump (rlhfbook.com)
| klelatti wrote:
| Web version with links, etc:
|
| https://rlhfbook.com/
| dang wrote:
| Thanks! We've switched to that above from
| https://arxiv.org/abs/2504.12501, and put the latter in the
| toptext.
| verdverm wrote:
| Last time I saw Nathan say something about the book, he's
| actively working on the next version and looking for feedback,
| check his socials
| leggerss wrote:
| You could say he's also learning from human feedback
| dang wrote:
| Related. Others?
|
| _RLHF Book_ - https://news.ycombinator.com/item?id=42902936 -
| Feb 2025 (37 comments)
___________________________________________________________________
(page generated 2026-02-07 23:00 UTC)