Picture the console of a modern warship, an air-defence node, or an armoured headquarters when the operator most needs it to be right. Dozens of automated alerts compete for one person's attention. Something must decide which alert that person sees first.

In almost every fielded system, that something is the anomaly score: the number that says how far a reading has strayed from its expected baseline. Ranking by that number feels objective, and it is easy to build. For the purpose that matters in operations, it is close to backwards. The fix is a design choice open to any force willing to make it.

The reason is simple. A statistical departure and an operational consequence are two different quantities. Systems that rank by the first stay silent about the second. An anomaly score measures how unusual a reading is. Mission impact measures what would happen if the condition behind that reading were real.

A large, attention-grabbing deviation in a subsystem that the mission does not depend on can matter far less than a small, quiet drift in a system that the mission rests on. Rank purely by the score, and the quiet, dangerous one sinks to the bottom of the queue. The loud, harmless one sits at the top wearing a red banner.

The operator's attention is the scarcest resource in the whole monitoring chain, and it gets spent exactly where it is needed least.

This failure mode is already written into official investigations. When the destroyer USS John S McCain collided with a tanker in the Singapore Strait in 2017, the United States National Transportation Safety Board found that steering control had been shifted between bridge stations without the helmsman realising it. The information that mattered most at that moment, which station actually held control, was in the system the whole time; it was never lost, only never absorbed.

The crew fought a phantom steering failure while the ship turned into traffic, and ten sailors died. Among the board's findings was that the design of the touch-screen steering and thrust control system itself increased the likelihood of the operator errors that led to the collision.

This is not an argument against anomaly scores. It is an argument for treating them as one input rather than the whole answer. A useful ranking lives on two axes, not one: how strong the statistical signal is and how much the affected function matters to the mission at that moment.

The four quadrants that result are not symmetric. A high-score anomaly in a low-impact system is usually noise or a minor fault that can wait. A low-score anomaly in a high-impact system is the dangerous case. It deserves a second look before it is dismissed, precisely because single-axis ranking would have thrown it away. Any ranking scheme worth having is built to protect that quadrant.

Cyber-resilience guidance from the United States National Institute of Standards and Technology defines the goal in mission terms: the capacity to anticipate, withstand, recover from, and adapt to adverse conditions. The point is to reduce the risk a mission carries by depending on cyber and computing resources.

That is a statement about consequences, not about statistical outliers, and it holds for any modern force, whatever flag it flies. A monitoring system that ranks by mission risk is doing what the resilience objective asks. A system that ranks by raw deviation is optimising a proxy and hoping it lines up with the mission, which it often does not.

The harder problem is that mission impact is not a fixed property that can be computed once and stored. It depends on what the affected subsystem contributes to the mission at that moment. That changes with phase, configuration, and tasking.

A degraded sensor that is irrelevant during transit can become critical during a terminal engagement. A communications path that is redundant in one posture becomes a single point of failure in another. A scheme that fixes its impact weights at design time will rank correctly in the phase it was tuned for and misrank in every other.

The impact model must move with the mission. Otherwise, it quietly becomes wrong at exactly the moments that matter most.

That need to change is also where these systems must be kept honest. The mapping from a subsystem to its mission consequence is a modelling choice. It should be made by the operators and engineers who know the platform. It should not be inferred silently by a model optimising whatever happened to be in its training data.

The basis for each impact estimate should travel with the alert. The operator can then see why something was ranked where it was and can override the ranking when the model's assumptions do not fit the situation in front of them.

The United States Government Accountability Office asks for the same discipline in its framework for accountable artificial intelligence: set the goals explicitly, document the basis, and check whether the scheme directs attention well rather than assuming it does. A ranking that cannot be inspected cannot be trusted. Nobody can catch it when a confident but wrong estimate sends scarce attention to the wrong place.

None of this requires exotic technology, only a design choice. Rank alerts by their consequence to the mission, not by the size of their statistical signal. Let the consequence model change as the mission changes. Keep it visible enough that a human can correct it.

The payoff is direct. In a contested environment, an operator will never have time to chase every flag. The cost of chasing the loudest alarm rather than the most dangerous one is measured in the events that were quietly important and got ignored. The systems that win the competition for human attention should be the ones that matter, not the ones that shout.

 

Still Interested?

Why not also read: Signature Management in Accelerated Warfare | Close Combat in the 21st Century by Dan Skinner.

Cove+ also has short learning courses on similar topics, including Introduction to Information Warfare.