Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
Saved in:
| Title: | Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. |
|---|---|
| Authors: | Dahlgren Lindström, Adam1, dali@cs.umu.se, Methnani, Leila1, leila.methnani@umu.se, Krause, Lea2, l.krause@vu.nl, Ericson, Petter1, pettter@cs.umu.se, de Rituerto de Troya, Íñigo Martínez3, i.m.d.r.detroya@tudelft.nl, Coelho Mollo, Dimitri4, dimitri.mollo@umu.se, Dobbe, Roel3, r.i.j.dobbe@tudelft.nl |
| Source: | Ethics & Information Technology; Jun2025, Vol. 27 Issue 2, p1-13, 13p |
| Database: | Applied Science & Technology Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: aci DbLabel: Applied Science & Technology Source An: 185750725 AccessLevel: 2 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AU" term="%22Dahlgren+Lindström%2C+Adam%22">Dahlgren Lindström, Adam</searchLink><relatesTo>1</relatesTo>, <i>dali@cs.umu.se</i><br /><searchLink fieldCode="AU" term="%22Methnani%2C+Leila%22">Methnani, Leila</searchLink><relatesTo>1</relatesTo>, <i>leila.methnani@umu.se</i><br /><searchLink fieldCode="AU" term="%22Krause%2C+Lea%22">Krause, Lea</searchLink><relatesTo>2</relatesTo>, <i>l.krause@vu.nl</i><br /><searchLink fieldCode="AU" term="%22Ericson%2C+Petter%22">Ericson, Petter</searchLink><relatesTo>1</relatesTo>, <i>pettter@cs.umu.se</i><br /><searchLink fieldCode="AU" term="%22de+Rituerto+de+Troya%2C+Íñigo+Martínez%22">de Rituerto de Troya, Íñigo Martínez</searchLink><relatesTo>3</relatesTo>, <i>i.m.d.r.detroya@tudelft.nl</i><br /><searchLink fieldCode="AU" term="%22Coelho+Mollo%2C+Dimitri%22">Coelho Mollo, Dimitri</searchLink><relatesTo>4</relatesTo>, <i>dimitri.mollo@umu.se</i><br /><searchLink fieldCode="AU" term="%22Dobbe%2C+Roel%22">Dobbe, Roel</searchLink><relatesTo>3</relatesTo>, <i>r.i.j.dobbe@tudelft.nl</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Ethics+%26+Information+Technology%22">Ethics & Information Technology</searchLink>; Jun2025, Vol. 27 Issue 2, p1-13, 13p |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=aci&AN=185750725 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10676-025-09837-2 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 13 StartPage: 1 Titles: – TitleFull: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Dahlgren Lindström, Adam – PersonEntity: Name: NameFull: Methnani, Leila – PersonEntity: Name: NameFull: Krause, Lea – PersonEntity: Name: NameFull: Ericson, Petter – PersonEntity: Name: NameFull: de Rituerto de Troya, Íñigo Martínez – PersonEntity: Name: NameFull: Coelho Mollo, Dimitri – PersonEntity: Name: NameFull: Dobbe, Roel IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 13881957 Numbering: – Type: volume Value: 27 – Type: issue Value: 2 Titles: – TitleFull: Ethics & Information Technology Type: main |
| ResultId | 1 |