Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.

Saved in:
Bibliographic Details
Title: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
Authors: Dahlgren Lindström, Adam1, dali@cs.umu.se, Methnani, Leila1, leila.methnani@umu.se, Krause, Lea2, l.krause@vu.nl, Ericson, Petter1, pettter@cs.umu.se, de Rituerto de Troya, Íñigo Martínez3, i.m.d.r.detroya@tudelft.nl, Coelho Mollo, Dimitri4, dimitri.mollo@umu.se, Dobbe, Roel3, r.i.j.dobbe@tudelft.nl
Source: Ethics & Information Technology; Jun2025, Vol. 27 Issue 2, p1-13, 13p
Database: Applied Science & Technology Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: aci
DbLabel: Applied Science & Technology Source
An: 185750725
AccessLevel: 2
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AU" term="%22Dahlgren+Lindström%2C+Adam%22">Dahlgren Lindström, Adam</searchLink><relatesTo>1</relatesTo>, <i>dali@cs.umu.se</i><br /><searchLink fieldCode="AU" term="%22Methnani%2C+Leila%22">Methnani, Leila</searchLink><relatesTo>1</relatesTo>, <i>leila.methnani@umu.se</i><br /><searchLink fieldCode="AU" term="%22Krause%2C+Lea%22">Krause, Lea</searchLink><relatesTo>2</relatesTo>, <i>l.krause@vu.nl</i><br /><searchLink fieldCode="AU" term="%22Ericson%2C+Petter%22">Ericson, Petter</searchLink><relatesTo>1</relatesTo>, <i>pettter@cs.umu.se</i><br /><searchLink fieldCode="AU" term="%22de+Rituerto+de+Troya%2C+Íñigo+Martínez%22">de Rituerto de Troya, Íñigo Martínez</searchLink><relatesTo>3</relatesTo>, <i>i.m.d.r.detroya@tudelft.nl</i><br /><searchLink fieldCode="AU" term="%22Coelho+Mollo%2C+Dimitri%22">Coelho Mollo, Dimitri</searchLink><relatesTo>4</relatesTo>, <i>dimitri.mollo@umu.se</i><br /><searchLink fieldCode="AU" term="%22Dobbe%2C+Roel%22">Dobbe, Roel</searchLink><relatesTo>3</relatesTo>, <i>r.i.j.dobbe@tudelft.nl</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Ethics+%26+Information+Technology%22">Ethics & Information Technology</searchLink>; Jun2025, Vol. 27 Issue 2, p1-13, 13p
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=aci&AN=185750725
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10676-025-09837-2
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 1
    Titles:
      – TitleFull: Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Dahlgren Lindström, Adam
      – PersonEntity:
          Name:
            NameFull: Methnani, Leila
      – PersonEntity:
          Name:
            NameFull: Krause, Lea
      – PersonEntity:
          Name:
            NameFull: Ericson, Petter
      – PersonEntity:
          Name:
            NameFull: de Rituerto de Troya, Íñigo Martínez
      – PersonEntity:
          Name:
            NameFull: Coelho Mollo, Dimitri
      – PersonEntity:
          Name:
            NameFull: Dobbe, Roel
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 13881957
          Numbering:
            – Type: volume
              Value: 27
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Ethics & Information Technology
              Type: main
ResultId 1