Dahlgren Lindström, A., Methnani, L., Krause, L., Ericson, P., de Rituerto de Troya, Í. M., Coelho Mollo, D., & Dobbe, R. (2025). Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. Ethics & Information Technology, 27(2), 1. https://doi.org/10.1007/s10676-025-09837-2
Chicago Style (17th ed.) CitationDahlgren Lindström, Adam, Leila Methnani, Lea Krause, Petter Ericson, Íñigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, and Roel Dobbe. "Helpful, Harmless, Honest? Sociotechnical Limits of AI Alignment and Safety Through Reinforcement Learning from Human Feedback." Ethics & Information Technology 27, no. 2 (2025): 1. https://doi.org/10.1007/s10676-025-09837-2.
MLA (9th ed.) CitationDahlgren Lindström, Adam, et al. "Helpful, Harmless, Honest? Sociotechnical Limits of AI Alignment and Safety Through Reinforcement Learning from Human Feedback." Ethics & Information Technology, vol. 27, no. 2, 2025, p. 1, https://doi.org/10.1007/s10676-025-09837-2.