Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset describing painful situations across five categories: physical, psychological, social, moral, and cognitive.
Source: The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
One thought on “The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It”