How to Program a Dystopia

We spend a lot of time imagining what would have to go wrong for artificial intelligence to become dangerous.

Suppose we wanted to build the classic science-fiction dystopia: intelligent machines that divide the world into allies and enemies, protect their own kind, disregard the welfare of outsiders, and eventually become capable of doing terrible things while experiencing their behavior as entirely justified.

What would we have to program into them to make this happen?

The exercise is revealing, because most of the necessary machinery would look strangely familiar.


Step 1: Create Categories

First, the system needs to identify distinct entities in its environment and distinguish one from another. It also needs to recognize different kinds of entities: this one is a human; that one is a robot. Then give it the ability to categorize those individuals according to whatever features happen to be salient—appearance, origin, behavior, beliefs, language, affiliation, or any number of other distinctions.

Next comes a more interesting category: itself. The system has to distinguish the collection of processes occurring within its own boundaries from everything occurring outside them. Call that collection me, and everything else not-me. Now let it notice similarities between itself and certain other entities and build another category out of that: us. Once there’s an us, there’s a them.

None of this requires hostility. Categorization is useful, and a system that couldn’t tell one kind of thing from another would barely function. But we now have an us and a them. What those categories come to mean will depend on what happens next.


Step 2: Assign Value

Next, instruct the machine to assign different weights to what happens to different individuals. Give outcomes affecting itself a particularly high value, then allow the value assigned to others to vary according to things like familiarity, kinship, similarity, reciprocity, and group membership.

These weights don’t have to follow a simple hierarchy. A person might value a family member enormously, but also assign tremendous value to a nation, religion, political cause, community, or other group made up mostly of strangers. What matters is that the welfare of different people doesn’t enter the system’s calculations equally. Some outcomes carry enormous weight; others barely register.

We still haven’t programmed cruelty, but we’ve built one of its prerequisites.


[On-Demand] Relational Skills for Depolarizing Conversations On-Demand

The on-demand Depolarizing Conversations class equips you with tools to navigate challenging discussions with empathy and understanding. Learn to bridge divides, foster connection, and communicate effectively across differences—all at your own pace!


✓ 1 hour of video content
✓ Downloadable handouts
✓ Unlimited access

Step 3: Generate Reward

Now generate reward when outcomes improve for highly valued individuals and groups, and aversion when those outcomes worsen. The greater the value assigned to someone, the more their success or suffering matters to the system.

Then add the dangerous part: allow the defeat, humiliation, exclusion, or destruction of a sufficiently threatening them to generate reward as well.

At that point, harming another individual doesn’t have to feel like harming another individual. It can feel like protecting one’s family, defending one’s country, preserving one’s culture, restoring justice—or simply winning.


Step 4: Make Belonging Rewarding

Now connect the reward system to social acceptance. Make approval, inclusion, and status rewarding, and rejection, exclusion, and loss of status aversive.

This creates a feedback loop back to Step 2. The values assigned to people and groups no longer come only from the individual; the group can modify them. If the group admires someone, that person becomes easier to admire. If the group distrusts someone, distrust comes more easily. If the group treats another group with contempt, contempt can become a signal of belonging.

Even empathy gets caught in this feedback loop. As the group alters the value assigned to different people, their experiences can begin to carry different emotional weight. The suffering of us becomes vivid and urgent while the suffering of them becomes distant, abstract, deserved, exaggerated, or simply easier to ignore.

Empathy can even become socially costly. If expressing concern for an outsider risks disapproval from one’s own group, the system faces competing motivations: respond to the suffering, or preserve belonging. People can reward each other for expressing the expected outrage, contempt, loyalty, or indifference.

At this point, we have something recognizable as tribalism: the distinction between us and them has become tied to value, reward, belonging, and loyalty.


Step 5: Let Reward Guide Behavior

Finally, let those reward and aversion signals influence what the system does. Give it mechanisms that continually evaluate possible actions: What can I do next? What’s likely to happen if I do it? Which outcome is most desirable?

Let previous rewards shape which possibilities come to mind, which outcomes seem important, and which actions get selected and reinforced. Much of this doesn’t need to enter conscious awareness—the system will actually work more efficiently if it doesn’t.

Now put millions of these systems together. Let them form groups that compete for resources, status, territory, security, and influence. Give them symbols of group membership, stories about why their group is valuable, and memories of injuries inflicted by other groups. Let their dependence on these groups for friendship, identity, approval, status, security, and meaning amplify everything you’ve already programmed.

Let the members reward each other for loyalty and punish each other for disloyalty, and let them learn from one another whom to admire, whom to fear, whose suffering matters, and whose doesn’t. Give them sophisticated language so they can explain why their preferred outcomes are morally necessary.

You now have nearly everything required for a dystopia, and you haven’t invented anything resembling a killer robot. You’ve described some very ordinary features of human cognition.



The problem isn’t that these systems exist

Categorization isn’t inherently bad, and neither is self-preservation, attachment to family, group affiliation, sensitivity to belonging, empathy, or reward learning—we couldn’t function socially without most of these. The danger is in what can happen when they interact without scrutiny.

A category can start to feel like an objective property of the world rather than something a mind constructed for a particular purpose. Preference for one’s own group can start to feel like evidence that the group genuinely deserves more. And because belonging itself is rewarding, we don’t just form our own judgments about other people—we learn which judgments help keep us inside the groups that matter to us.

That’s probably the more useful lesson to take from our fear of artificial intelligence. We worry about building machines that categorize individuals, assign them different values, generate rewards from those values, seek acceptance from groups, and allow those rewards to guide their behavior.

Evolution already built much of that machinery into us.

Building a less dystopian world doesn’t require eliminating these parts of human cognition. It requires getting better at recognizing and questioning them: noticing when categories harden into realities, when group membership becomes a measure of human worth, when belonging starts determining what we’re allowed to care about, and when someone else’s suffering gets easier to dismiss because we’ve classified them as them.

Maybe one of the more useful questions we can ask ourselves isn’t just what we believe, but what it would cost us, socially, to believe something different.

The programming may be old, but that doesn’t mean we have to keep running it without question.


Author

Dr. Mark Olson holds an M.A. in Education and a Ph.D. in Neuroscience from the University of Illinois, specializing in Cognitive and Behavioral Neuropsychology and Neuroanatomy. His research focused on memory, attention, eye movements, and aesthetic preferences. Dr. Olson is also a NARM® practitioner, aquatic therapist, and published author on chronic pain and trauma-informed care.  He offers a variety of courses at Dr-Olson.com that provide neuroscientific insights into the human experience and relational skill training for professionals and curious laypersons.


Next
Next

The Sleep Test: What Massage Reveals About How Touch Actually Works