Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You must be confused. This research is about reinforcement learning, not about large language models.


It's parroting human reinforcement.


It actually has no human data as input and learns by itself in the environment, that's the point of the accomplishment! :)


That's what humans do right? So it's parroting us.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: