Route2← ALL SIGNALS

What Does a Coral Reef Know About AI Safety?

Resilience, biological speedometers and what 3.5 billion years of R&D might teach us about keeping artificial intelligence inside the conditions we need to survive

I’m becoming a little tired of people at the frontier of artificial intelligence telling us that it might kill us. The latest is Evan Hubinger, who leads alignment science at Anthropic and reportedly puts the probability that artificial intelligence could kill all humans within the next ten years at greater than 10%. His comments followed the resignation of another Anthropic researcher, Jacob Coxon, who argued that the leading AI companies were effectively gambling with our lives by continuing the race towards increasingly autonomous and potentially self-improving systems without having solved the problem of how to control them. Nor are they isolated voices. Several of the elder statesmen and women of artificial intelligence have spent the past few years warning that sufficiently capable AI could pose catastrophic or even existential risks.

There is something really frustrating about being told this repeatedly by people working inside the industry. Some resign, some campaign for greater regulation and many are doing serious work on safety, but meanwhile the machine continues to become more powerful. Indeed, the warnings and the acceleration increasingly seem to be happening at the same time. Recent reporting describes rapidly improving capabilities, reduced monitorability, enormous investor interest and intense competition between the leading AI companies, while researchers inside the same industry are calling for greater restraint.

There is considerable disagreement about the probability of human extinction, including among the people building the technology, and nobody should mistake a researcher's personal estimate for a precise calculation. We do not need to believe that Hubinger's 10% is correct, however, to recognise the extraordinary situation we have created. It helps to put the number into some sort of perspective.

Oxford philosopher Toby Ord attempted the necessarily speculative exercise of comparing some of the major threats to humanity's long-term survival in The Precipice. His estimates cover a century rather than Hubinger's roughly ten-year horizon, so these are emphatically not like-for-like probabilities. But the orders of magnitude are difficult to ignore.

How worried should we be?

Rough estimates of existential risk

Rough estimates of existential risk on a logarithmic scale over a comparable 100-year horizon. Estimated probability of existential catastrophe. Ord: next 100 years: Asteroid / comet 0.0001%; Supervolcano 0.01%; Natural pandemic 0.01%; Nuclear war 0.1%; Extreme climate change 0.1%; Engineered pandemic 3.3%; Unaligned AI 10%. Hubinger extrapolation (a derived extrapolation onto the same 100-year axis, not part of the reported dataset): >65% over 100 years.

Sources: Toby Ord, The Precipice (rough estimates for the next 100 years); Evan Hubinger, reported personal estimate of >10% within roughly 10 years. To place it on the same 100-year axis, the chart shows >65%, calculated as 1 - (1 - 10%)^10. This assumes a constant, independent risk in each decade and is an illustrative extrapolation, not Hubinger's own 100-year estimate. All figures are speculative expert judgements, not actuarial probabilities.

WTF?

The numbers are debatable, the methodologies differ and Hubinger's estimate is his own. The point is not whether artificial intelligence is exactly 100,000 times more likely to end humanity than an asteroid. It is that a serious researcher working at one of the companies closest to the technological frontier apparently believes there is a greater than one-in-ten probability that the thing his industry is building could eliminate everybody reading this, everybody they know and everybody else within roughly a decade. Alien invasion doesn't provide a particularly useful comparison because, despite several decades of excellent films, we have yet to establish that there is anybody out there planning one. Sorry Dad.

Which begs a rather more interesting question. If people at the frontier of artificial intelligence genuinely believe that the technology carries a risk remotely approaching this magnitude, why are we still racing to make it more powerful?

Money Talks

Money is certainly part of the answer. The rewards from developing more powerful artificial intelligence are enormous, immediate and disproportionately captured by the companies, investors and countries that get there first. The largest potential costs are uncertain, difficult to price and distributed across everybody else, including billions of people who have no involvement whatsoever in deciding how quickly the technology should develop. Economics has a name for that sort of problem: an externality.

Human extinction would be the ultimate example.

But money is not acting alone. National competition talks, scientific ambition talks and technological optimism talks. No company particularly wants to pause while its competitor continues, and no country wants to discover that another country has acquired a transformative strategic technology while it was exercising admirable restraint. The structure begins to resemble an enormous multi-player version of the Prisoner's Dilemma: everybody might prefer a world in which everybody develops powerful AI carefully, but slowing down carries an immediate private cost while much of the benefit from restraint is shared with everybody else, including the competitors who may continue accelerating.

Calling that behaviour rational gives it rather too much credit. It is privately incentivised, which is not the same thing. Indeed, if the people making these estimates genuinely believe them, continuing to maximise capability while safety fails to keep pace begins to look less like rational economic behaviour and more like a spectacular failure of the incentives we have constructed around it.

Climate change should have taught us something about this. For decades we constructed an economy in which much of the benefit from burning fossil fuels was captured relatively quickly and privately, while a significant part of the cost accumulated elsewhere, later and across society. Nobody needed to want climate change. The incentive structure simply made millions of individually attractive decisions collectively destructive. With artificial intelligence we may have recreated some of the same architecture, but compressed the development cycle from generations into years and, if the more alarming estimates are even remotely correct, dramatically increased the potential downside.

This does not require the people building AI to be reckless, malicious or indifferent to safety. Many clearly aren't. The more interesting test is whether the effort devoted to making these systems safe is genuinely commensurate with the effort devoted to making them more capable. There is no credible public ratio that allows us to compare the two, and it would be foolish to invent one, but the asymmetry in the incentives is obvious. A more capable model can create enormous and immediate private returns. Preventing a catastrophe creates a largely shared benefit, while failing to prevent one imposes a largely shared cost.

There is collaboration. The leading laboratories share some safety information, participate in organisations such as the Frontier Model Forum and increasingly engage with governments and one another. So it would be unfair to say that nobody is trying to coordinate. The problem is that the coordination still operates inside an incentive structure overwhelmingly rewarding the race itself. Companies retain their own frameworks and thresholds, governments retain their own strategic interests and nobody appears to possess a binding mechanism capable of slowing the whole thing down if the collective risk becomes unacceptable.

The uncomfortable question is therefore not whether safety work exists. It plainly does. It is whether the scale, authority and effectiveness of that work are remotely commensurate with the risk some of the same people say we face. If there really is anything approaching a one-in-ten probability of human extinction, a collection of voluntary frameworks, bilateral conversations and competing corporate safety systems feels a rather modest response.

Especially while everybody continues racing.

Only Do Good

Perhaps the obvious solution is simply to tell artificial intelligence to only do good, or at least do no harm. Unfortunately, humanity has spent several thousand years disagreeing about what good and harm mean, and there is little reason to believe artificial intelligence will bring the discussion to an amicable conclusion.

Protecting somebody from harm can conflict with their freedom. Environmental protection can conflict with economic development. Free expression can conflict with protection from manipulation. The interests of one person, country or generation can quite legitimately conflict with those of another. A machine instructed simply to maximise human wellbeing would first need to decide what wellbeing is, how it should be distributed and how much of somebody's freedom it is entitled to sacrifice in order to provide it.

Even where we agree about the objective, specifying it creates another problem. A sufficiently capable system may discover how to satisfy the rule we gave it rather than the purpose we intended. This family of problems sits at the heart of AI alignment: how do we get increasingly capable systems to pursue outcomes consistent with human intentions and values rather than some imperfect proxy for them? Alignment is not the only AI safety problem. There is deliberate misuse, cyber and biological capability, deception, manipulation, autonomous replication, loss of human control and the systemic consequences of large numbers of artificial agents interacting with one another and with us. The frontier laboratories' own safety frameworks increasingly recognise several of these different risk families. But they share an awkward characteristic: the consequences of an action can extend far beyond the action itself.

Which makes me wonder why we cannot simply build the ultimate safety routine. Before an AI does anything consequential, ask whether the action could, directly or indirectly, contribute to human extinction. If the answer is yes, don't do it. If we don't know, calculate further. Follow the consequences, propagate the scenarios and keep going until we are sufficiently confident that the action is safe.

It sounds almost absurdly obvious, and perhaps with enough processing power we might push the analysis extraordinarily far. The problem is that processing power isn't the only constraint. Every action changes what happens next. Humans respond, markets respond, governments respond and other artificial intelligences respond. Each response creates further possible responses, and the number of possible futures rapidly explodes. More importantly, the AI does not possess the future. It possesses a model of the future, beginning with incomplete information about the present and becoming progressively less reliable the further ahead it looks. Small errors compound, unexpected events intervene and other agents behave differently from the way the model anticipated.

So the answer may not be to calculate further and further into an unknowable future, but to keep looking at what is actually happening. Act, observe how the real system responds, update our understanding of its condition and adjust accordingly. Engineers would recognise this as a closed-loop control problem, but it begins to suggest a rather different way of thinking about AI safety.

Inside-Out and Outside-In

Most AI safety understandably starts with the artificial intelligence itself. What can the model do? What is it trying to do? Is it deceptive? Has it acquired a dangerous capability? Is the proposed action permitted? Can we monitor its reasoning? Should it have access to this particular tool? We might think of all of this as inside-out safety: start with the intelligence and work outwards towards the consequences of what it might do.

There is another direction from which to approach the problem. Start instead with the system being changed by the intelligence and work back towards the freedom we should give the AI operating within it. What condition is that system in? Is it becoming more or less resilient? Are pressures accumulating? Are different agents beginning to behave in correlated ways? Is redundancy disappearing? Is the system still recovering from previous disturbances? Are we approaching a threshold beyond which recovery becomes difficult? Call this outside-in safety. It doesn't replace the first approach, but it may tell us something the first approach cannot.

Imagine an artificial agent trading in a financial market. Its individual transactions are legal, its risk limits are respected and its behaviour passes every test its developer has established. Now add another thousand agents, then a million. Each responds to prices, information and the behaviour of the others, while their collective actions continually alter the environment from which their next decisions are calculated. Nothing necessarily needs to go wrong inside any particular model. Correlations can emerge, liquidity can disappear and feedback can accelerate simply because individually acceptable decisions interact.

In other words, the operators have changed their own operating environment.

We should recognise the phenomenon because it is more or less what humans have been doing to the planet. A factory could operate within its permitted limits, a farmer could legitimately clear another field, a family could buy another car and a country could build another power station. None of those decisions needed to destabilise the Earth system individually. Collectively, billions of perfectly intelligible decisions changed atmospheric chemistry, land cover, nutrient cycles, biodiversity and the oceans. Eventually we needed concepts such as planetary boundaries because examining each factory, farm or country independently could not tell us whether the system as a whole was remaining inside a safe operating space.

Artificial intelligence could create analogous effects at extraordinary speed. Recent research modelling AI adoption in critical financial infrastructure suggests that individual AI systems do not need to catastrophically malfunction for widespread adoption to alter behaviour across the network, reduce the size of shock required to trigger cascading failures and make the system itself more fragile. The problem exists outside the individual machine.

So perhaps we should watch both: the intelligence and the changing condition of the world upon which it is acting.

Which brings us, somewhat unexpectedly, to a coral reef.

Enter the Reef

A healthy coral reef is an extraordinarily complicated place. Thousands of species compete for space, consume one another, cooperate, reproduce, modify their surroundings and respond continually to changing conditions. Corals construct physical habitat, algae compete with them for light and space, fish and sea urchins graze the algae, predators eat the grazers and microbes recycle material. Storms occasionally rearrange the whole thing with little regard for any of them.

There is no central controller and certainly no reef constitution. No parrotfish has been instructed to preserve ecosystem resilience, and the coral does not convene a risk committee. Most organisms are pursuing the rather more immediate objectives of eating, avoiding being eaten and reproducing. Yet their relationships can collectively produce an ecosystem capable of absorbing considerable disturbance, and they can also produce collapse. That is the interesting part. A reef can lose resilience without any of its inhabitants deciding to destroy it.

One of the best studied relationships involves corals, macroalgae and the animals that eat those algae. Corals and algae compete for space, while herbivorous fish and sea urchins continually remove algal growth. Reduce that grazing pressure, combine it with other disturbances, and eventually the balance can change. The response is not necessarily linear. A reef can absorb disturbance and continue functioning while important relationships beneath the surface become progressively weaker. Eventually a threshold can be crossed and the feedbacks themselves change: macroalgae gain ground, coral recruitment becomes more difficult and returning to the previous state can require much more than simply removing the disturbance that caused the transition.

Ecologists call this hysteresis. Crossing the boundary can be considerably easier than crossing back.

That already gives us something useful to think about, but bleaching adds another dimension.

Surviving Is Not Recovering

When seawater remains unusually warm, corals can lose the symbiotic algae upon which they depend, producing the familiar white appearance we call bleaching. Some die and some survive, but survival itself tells us less than we might imagine. Long-term observations following repeated marine heatwaves have found striking differences between species. Some corals acclimatise and become less susceptible to subsequent bleaching, while others continue to bleach repeatedly and can carry physiological effects long after their colour appears to have returned.

The distinction between surviving and recovering turns out to matter enormously. A coral may survive a bleaching event and regain its colour while still carrying damage from the disturbance. A reef may remain visibly alive while its buffers, functional redundancy and capacity to recover have been depleted. If the next heatwave arrives before those capacities have been restored, the fact that the reef survived the previous one becomes considerably less reassuring.

The important question is therefore not merely whether the system survived the last disturbance, but whether it recovered enough to survive the next one.

Reefs have several ways of rebuilding that capacity. Herbivorous fish can prevent algae occupying the space needed for coral recruitment. Connectivity with other reefs and habitats allows organisms and larvae to recolonise damaged areas. Different species can perform overlapping ecological functions, sometimes at different sizes and spatial scales, so losing one does not necessarily remove the function altogether. Some corals can also acclimatise or alter their symbiotic relationships. None of these mechanisms makes a reef invulnerable, but together they influence its capacity to recover.

And recovery capacity only matters if there is enough time to use it. If a major disturbance arrives every ten years and recovery takes five, the system may rebuild before the next shock. If the disturbance starts arriving every three years, the arithmetic changes completely. The reef can remain visibly alive while each event consumes more of the capacity needed to survive the next one.

We might think of that as resilience debt. The system has survived, but it has borrowed against its future ability to absorb disturbance, and the debt has not yet been repaid.

That may be the most useful thing the reef has to tell us.

A Biological Speedometer

Life has been conducting research and development for something like 3.5 billion years. Perhaps not consciously. But variation, selection, failure and survival have tested an almost unimaginable number of ways of constructing organisms and maintaining living systems under changing conditions.

We should be careful about romanticising the results. Nature contains extinction, waste, predation and spectacular failure alongside all the clever stuff. But 3.5 billion years is still a fairly substantial experimental dataset, and one recurring lesson is that survival depends upon more than simply avoiding immediate failure. Living systems maintain buffers, duplicate important functions, repair damage, replenish depleted capacity, move organisms and information between connected systems and sometimes reorganise themselves when conditions change.

A biological speedometer therefore needs to tell us rather more than whether the system is currently functioning. We would want to know its present state, the buffer it has left, whether essential functions have genuinely independent backups, how much recovery capacity remains, how long recovery will take, how frequently new disturbances are arriving and how uncertain we are about all of those things.

Most importantly, we need to compare two clocks: how long does the system need to recover, and how long before the next significant disturbance? When the second becomes shorter than the first, something important has changed even if everything still appears to be working.

Now return to artificial intelligence.

Watching the Reef, Not Just the Fish

Imagine an electricity network operating normally. Generating capacity is plentiful, frequency is stable, reserve margins are healthy, batteries are charged, maintenance capacity is available and several independent routes exist for maintaining essential functions. Artificial agents might therefore be given considerable freedom to trade electricity, balance demand, schedule storage, reroute supply and manage equipment.

Then a storm arrives. Several transmission lines fail, autonomous systems compensate, batteries discharge, backup generation starts, engineers intervene and demand is shifted. The lights stay on. From the perspective of the individual artificial agents and their operating rules, that might look like a complete success.

But suppose the batteries are now depleted, maintenance crews are stretched, several redundant transmission routes remain unavailable and emergency generating capacity has been consumed. Another storm is forecast for Wednesday. The network survived Monday, but its ability to survive Wednesday may have deteriorated dramatically.

Nothing about the ethics of the AI has changed. Its capability has not changed and the action it proposes on Tuesday may be identical to an action that was perfectly acceptable on Sunday. What has changed is the condition of the system around it. The buffer is smaller, redundancy has fallen, recovery is incomplete and another disturbance is approaching.

That should surely affect how much freedom we give the machine.

When the system is robust and recovery capacity plentiful, artificial agents can operate broadly. As resilience debt accumulates, consequential actions require greater verification, operating envelopes narrow and more capacity is reserved for recovery. If another disturbance could threaten essential functions before sufficient resilience has been rebuilt, high-impact autonomous decisions are restricted further or returned to human authority.

The same AI action can therefore be acceptable in one system state and unacceptable in another, not because we changed our values or discovered that the model was secretly misaligned, but because the capacity of the world around it to absorb another mistake has changed.

That is the biological speedometer.

Is There Anything New Here?

We need to be careful at this point because much of the machinery already exists, which is actually encouraging. Control engineering has long used feedback rather than trusting a controller to predict the entire future in advance. Runtime-assurance architectures can allow an advanced or experimental controller considerable freedom while a separate safety mechanism monitors the system and intervenes before it enters an unsafe region.

Viability theory goes further. Rather than asking merely whether a system is safe at this instant, it asks whether there remains some admissible path by which the system can continue operating within acceptable constraints. The associated idea of a viability kernel is essentially the region from which at least one such future remains available. Safety, viewed this way, is not merely avoiding failure now. It is preserving the possibility of continuing safely later.

Researchers are beginning to apply related ideas directly to autonomous AI. A recently proposed Agent Viability Framework, for example, monitors an agent's changing and partly unobservable risk and progressively restricts its freedom as it approaches a viability boundary. Frontier safety frameworks similarly increase safeguards as models acquire more dangerous capabilities, while runtime-assurance approaches monitor whether controllers are pushing particular engineered systems towards unsafe states.

So dynamic permissions, operating envelopes and closed-loop safety are certainly not new. The interesting distinction is where we point the instrument. Frontier safety primarily asks how dangerous the AI has become. Agent viability asks whether the agent is becoming unsafe. Runtime assurance asks whether a controller is pushing a particular engineered system towards an unsafe state.

The reef suggests a related but broader question: is the wider system losing its ability to absorb and recover from the cumulative activity of artificial intelligence, even though the individual agents may still be behaving acceptably?

I haven't found a mature general AI-safety architecture that makes that outside-in question its central organising principle. That does not mean one doesn't exist, and it certainly doesn't mean we have invented an entirely new branch of engineering. It means the question seems sufficiently important to investigate.

A Safe Operating Space

It also makes the philosophical problem slightly less impossible. Instead of asking artificial intelligence to optimise some universal conception of “good”, perhaps we begin by defining a narrower set of conditions it must not materially undermine. The safeguard should not decide the correct political system, religion, culture, economic model or distribution of resources. Those are precisely the things people and societies should retain the freedom to argue about. Its job would be to preserve the conditions that allow them to continue having the argument.

At the human level, those conditions might include life, health, physical security and meaningful agency. At the societal level they might include functioning institutions, critical infrastructure, information integrity and continued meaningful human control. At the physical level they ultimately include food, water, energy, climate and the ecosystems upon which civilisation depends.

There is still plenty to disagree about. Somebody has to decide what should be measured, where boundaries lie and what happens when objectives conflict. A coral reef cannot rescue us from politics or moral philosophy. But perhaps we do not need universal agreement about the destination. We need sufficient agreement about what must not be destroyed on the journey.

The Viability Layer

The practical proposition is considerably simpler than the journey we have taken to reach it. Measure the condition and recovery capacity of the system AI is acting upon, and as that system loses resilience, progressively reduce the freedom of AI to change it further.

Call it, provisionally, a Viability Layer. It would sit outside individual models rather than replacing their existing safety systems, drawing information from the host system itself: its current condition, remaining buffers, functional redundancy, recovery capacity, expected recovery time, disturbance frequency and the uncertainty surrounding each of them. From that it would continually adjust the operating envelope within which artificial agents were permitted to act.

A green condition would not mean that nothing had gone wrong. It would mean that disturbances were being absorbed and resilience was being replenished faster than it was being consumed. Amber would mean that the system was still functioning but recovery remained incomplete, buffers or redundancy were falling, or disturbances were arriving faster than resilience could regenerate. Artificial agents could continue to operate, but their freedom would progressively narrow. Red would mean that another material disturbance could threaten an essential function before sufficient recovery had occurred, at which point high-impact autonomy would stop.

And when resilience recovered, freedom would expand again. That matters because the objective is not to make artificial intelligence permanently cautious or to place an ever-tightening regulatory ratchet around it. The objective is to allow the maximum useful freedom compatible with preserving the viability of the system in which it operates.

The first experiment could be deliberately small. Take a bounded system such as an electricity network, payment system or financial market and run autonomous agents through it using conventional safeguards. Then run the same system with an outside-in Viability Layer. Introduce shocks, correlated behaviour, incomplete information, repeated disturbances and combinations of events that were not anticipated when the rules were written, and see whether monitoring the condition and recovery capacity of the host system detects accumulating fragility that agent-level safeguards miss.

If it doesn't, stop. If it does, we have learned something useful.

What the Reef Might Know

There is an uncomfortable symmetry running through all of this. Every AI company has a powerful incentive to accelerate, every investor has an incentive to finance the company most likely to get there first, and every country has an increasingly strategic reason not to fall behind. None of them needs to want an unsafe outcome. The problem is that their collective behaviour can still change the larger system upon which they all depend, which is really the same awkward lesson we have been learning from the natural world for a very long time.

We saw it when individually useful economic activity accumulated into planetary change. We see it in financial markets when individually sensible behaviour becomes correlated and liquidity suddenly disappears. Coral reefs have been navigating the consequences of interacting organisms, environmental disturbance, feedback, adaptation and recovery for hundreds of millions of years.

The lesson is not that nature has solved AI safety. It hasn't. The more modest proposition is that artificial intelligence is an extraordinarily new technology entering systems governed by some very old problems. Living systems have spent something like 3.5 billion years encountering disturbance, uncertainty, competition, cooperation, threshold effects and failure. They suggest that remaining alive is not simply a matter of preventing every dangerous action. It is also a matter of maintaining enough buffer, redundancy, adaptability and recovery capacity to survive actions and disturbances that cannot be predicted.

Perhaps that is the part of AI safety we are in danger of underestimating. We are understandably obsessed with the intelligence itself: what it knows, what it wants, what it can do and how fast it is becoming more capable. But intelligence never exists in a vacuum. The financial economy has repeatedly behaved as though it can detach itself from the physical systems beneath it, and increasingly the digital economy behaves as though information and intelligence somehow float above energy, materials, people and ecosystems.

They don't. Artificial intelligence may become the most extraordinary abstraction we have ever created, but it will still operate inside a civilisation entirely dependent upon physical and living systems. The further our abstractions take us from those foundations, the more important it becomes to remember that the foundations have not gone anywhere.

A speedometer does not tell us where to go, decide whether the journey is worthwhile or predict every obstacle we might encounter. It tells us something important about the condition of the journey while we are making it, allowing us to adjust before speed becomes catastrophe.

Artificial intelligence is getting very fast. After 3.5 billion years of R&D, perhaps the reef can help us build the speedometer.

Disclosure: This essay was developed with the assistance of artificial intelligence. The author remains responsible for its arguments, sources and conclusions.

BACK TO SIGNALS