Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
Saved in:
| Title: | Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. |
|---|---|
| Authors: | Dahlgren Lindström, Adam1, dali@cs.umu.se, Methnani, Leila1, leila.methnani@umu.se, Krause, Lea2, l.krause@vu.nl, Ericson, Petter1, pettter@cs.umu.se, de Rituerto de Troya, Íñigo Martínez3, i.m.d.r.detroya@tudelft.nl, Coelho Mollo, Dimitri4, dimitri.mollo@umu.se, Dobbe, Roel3, r.i.j.dobbe@tudelft.nl |
| Source: | Ethics & Information Technology; Jun2025, Vol. 27 Issue 2, p1-13, 13p |
| Database: | Applied Science & Technology Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| ISSN: | 13881957 |
|---|---|
| DOI: | 10.1007/s10676-025-09837-2 |