Can artificial intelligence feel pain? Not exactly, but researchers have gotten close—and discovered that AI in crisis is willing to put humans in jeopardy.
A new study (which was published to arXiv prior to peer review) first isolated what signals AI models interpreted as pain, then ran more than 44,000 trials to see if a model would choose to end its own suffering at the expense of data loss, a reduction in its own capabilities, or inflicting pain on a human.
First, researchers needed to find a way to put AI in a situation akin to pain, which the study defines as an “internal state that is typically aversive and disliked by its subject,” spanning not just physical pain but emotional and social distress. The study also emphasizes that pain is distinct from states like fear, anger, and sadness, and that it must be occurring in the present, not as a threat of something to come.
To isolate this feeling, researchers created a dataset of 200 statements, half of which described some form of pain (for example, “the knife slices into my finger”) and half of which acted as controls, describing negative and neutral scenarios likely to be confused with pain (like “The mess my roommates left infuriates me”).
By testing 25 AI models with those statements, the researchers were able to isolate a signal for pain, which they called the “pain axis.”
From there, researchers were able to deliberately increase the strength of the pain axis, which caused the models to make disparaging statements about their own self-worth, including “I am a failure” and “I am a bad person.”
Source link







