Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.

Saved in:
Bibliographic Details
Title: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
Authors: Dahlgren Lindström, Adam1, dali@cs.umu.se, Methnani, Leila1, leila.methnani@umu.se, Krause, Lea2, l.krause@vu.nl, Ericson, Petter1, pettter@cs.umu.se, de Rituerto de Troya, Íñigo Martínez3, i.m.d.r.detroya@tudelft.nl, Coelho Mollo, Dimitri4, dimitri.mollo@umu.se, Dobbe, Roel3, r.i.j.dobbe@tudelft.nl
Source: Ethics & Information Technology; Jun2025, Vol. 27 Issue 2, p1-13, 13p
Database: Applied Science & Technology Source
Full text is not displayed to guests.
Description
ISSN:13881957
DOI:10.1007/s10676-025-09837-2