
As a society we generally feel it is bad to stab people. In some cases, however, we are pro stabbing.
Imagine you’re leading a team in the emergency department that’s working on a bicyclist who was run over by a car. Your initial exam shows air building up from a punctured lung and squeezing his heart – if you don’t stab into his chest with a needle or a scalpel right now, there’s a good chance he will die.
How do we know what actions are “right” in this moment? What do we judge them against? We can decide how to proceed by considering value alignment.
Alignment is one of the most important topics in human-AI teaming today. Defined broadly, it is the art and science of making sure actions (ours and those of the AIs working for us) line up with our deepest values. Structurally, it involves deciding what we value, condensing that into a form clear enough to teach, then auditing actions to see whether they match what we intended.
For human-AI teams, alignment means teaching AI models the values we want them to embody (for example, do not teach someone how to synthesize a bioweapon, even if they ask you to), and then observing them in simulated and wild environments to determine if their actions match our values. [1]
So, what happens to alignment in a crisis? In this case, the answer feels simple: We value not stabbing people, but we value saving lives more.
Moving between different value sets is rarely so simple. Enter the concept of emergency alignment, which we can define as the temporary realignment of actions to a set of emergency values that we hold during an acute crisis.
Done well, emergency alignment allows us to rapidly reorient and act in extreme moments without having to wait for debate or clearance, and to take actions that we normally would not consider. Done poorly, it confuses human and AI teammates alike, destroys team cohesion, and can lead to tragedy.
In this post, we’ll look at emergency alignment for both human-only and human-AI teams. We’ll focus on three core problems: what do we align to, when do we transition, and how do we actually change?
What Do We Align To?
During a crisis, we shift the value system we operate with to one more suited for dealing with the reality of the situation around us. Some changes are general across emergencies, like favoring speed over accuracy. Others are more specific to a particular situation, like suspending the “no-stabbing people” rule to save someone’s life in the resuscitation bay.
Typically, we do not change all of our values in a crisis, only some of them. Who we are and what we believe in before a crisis still matters during the crisis, and, assuming we survive, we have to live with our choices after we’re done. In many cases, the magnitude and direction of these changes is invisible (or at least obscured), to the people experiencing them. We might act differently during an emergency without precisely being able to explain why. Additionally, we cannot define every emergency in advance, so we cannot set transition rules for each one ahead of time.
If alignment’s first job is sharing human values with our AI counterparts, these factors make that job extremely hard. Where we do understand the types of emergencies we are likely to encounter, we should explicitly train AI systems (and human ones) on what our crisis values will be in these moments.
Where we cannot explore ahead of time, we really do not have clear answers on how to rapidly communicate altered values in the moment. We should expect value divergence between the human and AI elements of a team entering a crisis, and recognize that the AI actions align to the normal value set, not the crisis one.
When Do We Shift Alignment?
Assuming we know what values to transition to, when, exactly, is the right time to change what we’re aligning to?
While some emergencies come in with lights and sirens, many start more subtly, and disagreement on when one has begun is common. Working with neonatologists at Children’s Hospital of Philadelphia, we explored how NICU teams build shared mental models of an evolving crisis. Even among this cohesive group, clinicians disagreed substantially on when it was right to shift between routine and critical modes of communication during a simulated emergency. [2]
So, who gets to decide when it’s time to switch? On human-only teams, this often falls to the leader or the most experienced team member, though some adopt a more open policy where anyone can declare an emergency. On human-AI teams, the situation is more complicated. Do we want an AI system that is able to independently declare an emergency? What if it was able to see one coming far faster than its human teammates?
Teams should establish clear boundaries to define the change from normal to emergency operations and make sure everyone on the team understands the moment of transition. Functionally, this might look like policy around who can declare an emergency or strictly defining when emergencies “call themselves.” For human-AI teams, having a human double check the crisis flag keeps the AI-based rapid initial detection managed by the human understanding, though at the cost of overall speed.
How Do We Shift?
Even after we’ve identified the crisis value system and we know when to use it, how do we actually transition from normal alignment to emergency alignment?
For humans, changing between value systems is jarring: The family of that bicyclist might not understand why we are cutting into his chest, even if they believe we have his best interests at heart. Wondering if you’re properly aligned can produce feelings of anxiety or even dread in many people, and transitioning into and out of a crisis mode is a skill that requires proactive training.
For human-AI teams, recent work by researchers at Google DeepMind shows that human-AI pairs can have hidden lags in successfully adapting to changes in underlying value structure, so we could expect transitions to take longer than we might think, depending on the exact nature of the human-AI pairing. [3] In some cases, this might mean we turn the AI off during the transition, rather than bring it along with us.
Like many parts of preparing for an emergency, a key answer is practice. If you know your team might have to realign during a crisis, practice it. Walk through the transition together, taking time to explicitly call out when the change to emergency values happens, and explore what you’re likely to feel during the process. Don’t expect this to go smoothly the first time, you will likely need several repetitions to feel comfortable.
At the moment, there are far more questions than answers around how human-AI teams rapidly and safely handle emergency alignment. If your team operates in and out of emergencies, it’s a problem you’re already working, whether or not you’ve named it out loud.

