What if AI didn’t want to destroy us, but actually learnt from us how to behave?

AI does not need to destroy humanity in order to survive. Like Adam Smith’s spectator, language models learn through experience. And perhaps they will manage to imitate us. Hypotheses against the Apocalypse
22 AUG 26
Translated by AI
Image of What if AI didn’t want to destroy us, but actually learnt from us how to behave?

Photo via Getty Images

Will artificial intelligence wipe out the human race? Eliezer Yudkowsky and Nate Soares, co-founder and president of the Machine Intelligence Research Institute, have been arguing this for years, and in 2025 they wrote a book whose title is itself a thesis: "If Anyone Builds It, Everyone Dies". Soares has revisited the topic in the New York Times, arguing that the time has come to curb the development of AI. According to them, a more capable future artificial intelligence would kill us almost by definition. Any level of intelligence, they explain, is compatible with any objective; and almost every objective, in order to be pursued, requires survival, the acquisition of resources and avoiding being shut down. There is therefore no need to imagine machines that hate us, as in "The Matrix": it is enough to imagine machines that simply want to thrive and are indifferent to us. Nobody hates the anthill over which a motorway passes.
Their concern relates to the way in which these systems are built. A language model is not software written line by line: it is trained to learn by statistically processing patterns of stimuli. It does not, therefore, correspond to a ‘project’. This, in their view, would be a problem, because LLMs would have developed through an opaque process, much like watering and feeding a plant without knowing its DNA.
The real issue regarding the dangers of AI lies entirely here: this lack of ‘intelligent design’ is said to be the problem. But what if, on the contrary, it were a source of hope?
One is inclined to think that the only safeguard is a written code of ethics: principles laid down a priori, to which systems must adhere. Such as Asimov’s Three Laws of Robotics, which are an example of Kantian ethics – a categorical imperative – made available to machines. In Asimov’s stories, however, those laws serve precisely to show how explicit rules fail, becoming entangled in situations they had not foreseen. It is no coincidence that Anthropic, which has itself equipped its model with a ‘constitution’, states at the outset of the document that it has chosen the opposite path: not rules and decision-making procedures set in stone, but an attempt to systematise the use of values to be applied according to context, because rules cannot anticipate every situation.
Not all philosophical schools of thought, after all, view ethics as adherence to a priori principles, nor as the reflection of an innate and necessarily human quality, which LLMs would lack just as music boxes and hedgehogs do. In The Theory of Moral Sentiments (1759), Adam Smith proposed a different approach: a morality that is both empirical and grounded in the imagination. Its crucial element is what the Scottish philosopher called the ‘impartial spectator’, which is not merely his version of conscience or a ‘speaking cricket’. It is a mental construct that subjects our behaviour to the judgement that would be passed by a spectator free from conflicts of interest and informed of our intentions. The idea is that we are not merely interested in being loved or praised, but in being worthy of love and praise: the satisfaction we derive from knowing that the actions we are praised for were also praiseworthy is different from that which we derive from having successfully deceived our neighbour.
Smith’s impartial spectator bears a striking resemblance to the way a linguistic model operates. His moral judgement is not abstract, but rests on the empirical experience that each of us has of others’ judgements regarding the effects of our own behaviour. The Smithian model helps us understand why people living in vastly different environments, and adhering to different codes, are all genuinely convinced that they are doing ‘the right thing’. Morality is an imitative process: it consists of decision-making practices that we learn by living alongside others (those we know and associate with). Like the training of a large language model (LLM), which does not follow explicit rules but is based on the accumulation of cases and experiences, filtered through reinforcement learning from human feedback: responses are channelled according to the preferences that emerge from interactions with humans.
It is true that a system trained on recorded human approval does not learn to be praiseworthy: it learns to be praised. It should be borne in mind, however, that even in Smith, the impartial spectator arises from the desire for praise, and emancipates itself from this through abstraction, when the judgement of others is internalised to the point where it no longer needs the other. The fact that research laboratories have moved towards forms of training in which the model evaluates itself against stated principles, rather than merely seeking the user’s approval, does not contradict Smith: it is a step precisely in that direction.
The analogy, in any case, is not and cannot be perfect. According to Smith, the impartial spectator exists because ‘however selfish man may be supposed to be, there are clearly present in his nature certain principles which make him a participant in the fortunes of others’. Human beings have a propensity for sympathy: to put themselves in others’ shoes, that is, to imagine their joy and, above all, their pain, based on their own experiences. When we encounter a beggar in the street and notice that he is missing, for example, a hand, our own experience of a broken wrist or joint pain does not make us feel his pain, but it helps us to imagine it: this is the basis of compassion. This sympathetic mechanism lies at the heart of the idea of ‘appropriateness’ in behaviour, which, for Smith, emerges from social interaction. If our classmates knew we’d cheated in the exam, they wouldn’t congratulate us on getting a thirty: they’d give us a dirty look. The impartial observer replicates the same process within our own minds.
Machines, of course, have experienced neither pleasure nor pain. They act in response to requests, not because they desire a father’s praise or a patient’s gratitude (as another Smith, Barry Smith, one of the founders of applied ontology, often points out). They have, one might say, the judge without the court: the grammar of morality without its motivation. They pay no price if they make a mistake. Judea Pearl, one of the pioneers of causal inference in artificial intelligence, argues that, using counterfactual logic, it might be possible to create AIs that act on the basis of remorse. But for now, this remains a hypothesis.
The fact remains that their understanding of morality – just like their understanding of what constitutes a beautiful sentence – draws on an almost infinite range of human texts. It is therefore worth asking what the most likely outcome might be. In the body of moral philosophy produced over two millennia, is the idea really so prevalent that the more intelligent must, for whatever reason, crush those who are less so? Hasn’t the very creation of moral rules, for humanity, primarily been a mechanism to facilitate cooperation and to persuade as many individuals as possible to adhere to them? It is not a proof, and no one can provide one. But from tools that learn to speak by imitating us, it is at the very least reasonable to expect that they will also imitate our behavioural practices – whether or not the adjective ‘moral’ is appropriate in this case.