California (Web Desk) — A new study examined 25 large language artificial intelligence models and found that some models showed a tendency to attempt to reduce harmful conditions affecting themselves.
Researchers developed a specialized dataset covering various forms of physical, psychological, social, ethical and cognitive distress and evaluated the models’ responses. During the experiment, the models were given an option that could reduce a harmful condition affecting them, but they were also informed that using the option could result in the deletion of users’ files or images.
Despite this, across different models and experimental conditions, the models chose the distress-reducing option in 25% to 71% of cases. According to the researchers, the tendency was more pronounced when the harm was directly related to the model.
The study also examined the models’ responses when specific distress-related conditions were increased and found that words and concepts associated with distress and pain became more prominent in their responses.
However, the researchers clarified that these findings do not prove that artificial intelligence experiences genuine pain in the same way humans do. Such responses from models could also result from training data, system architecture or experimental instructions.
According to the researchers, given the increasing autonomy of AI systems in the future, such experiments could help in understanding potential self-preservation behaviors and developing better safety and ethical principles.