Resilience over Perfection

Picture the boardroom. The deck is loaded, the clock is ticking, and the chair leans forward with that familiar question: “So… are we secure?” The slide behind you is full of colors, domains, and scores that look reassuring from a distance, but you know none of it adds up to a clean yes. In a world of shifting software supply chains, cloud entanglements, and third-party dependencies, “secure” is no longer a destination. It is a moving weather system. This narration is drawn from the Wednesday “Headline” feature in Bare Metal Cyber Magazine, developed by Bare Metal Cyber, and it is about changing that conversation so it finally matches the world you actually operate in.

The basic problem is not that your program is weak or that your team is unaware of the risks. The problem is the story the organization expects you to tell. Boards, regulators, and business leaders are wired to look for finish lines. They ask when things will be “done,” when the risk will be “under control,” when you will “get to green.” Over the next few years, that expectation becomes less and less compatible with how technology really behaves. The more digital the business becomes, the more any honest security story has to admit that there is no stable end state. That is not failure. That is physics.

If you sit in a leadership role long enough, you feel the gravitational pull of “secure enough.” It shows up in budget cycles, where funding is pitched as the last big push to close the remaining gaps. It shows up in hallway questions, when someone asks how close you are to the target state. It shows up in your own internal language, when you talk about “hardening” environments and “closing” workstreams as if they will stay closed. None of this is malicious. It is how human beings try to simplify complexity so they can move on to the next problem.

The mirage is reinforced by the way organizations frame accountability. Regulators and auditors often push toward binary answers: compliant or non-compliant, gap open or gap closed. Insurers ask for checkboxes. Internal assurance functions like clean statements and clear dates. In response, security teams smooth out the rough edges of reality. You promise to close a control gap by a quarter end. You talk about “achieving” a target maturity level. You present heatmaps that imply a stable landscape. Over time, those simplifications harden into promises that nobody could actually keep in a dynamic environment.

Living inside that “secure enough” mirage has real costs. It pushes you toward cosmetic wins that photograph well in a board pack but barely move the organization’s ability to withstand a serious incident. It quietly discourages people from speaking honestly about residual risk, because every admission of uncertainty sounds like an admission of failure. It encourages leaders to view every new control or tool as a step closer to completion, rather than a living obligation to operate, monitor, and adjust that control as the business changes. The first step toward resilience is to stop playing along with that illusion and say out loud that “never done” is not an indictment of your program; it is the nature of the terrain.

This is where maturity models enter the story. On paper, capability and maturity frameworks are navigation aids. They help you describe which controls exist, how consistently they are operated, and where you might invest next. In many organizations, though, they slip into theater. The model turns into a large spreadsheet of domains and subdomains, each sitting on a one-to-five scale that is, at best, loosely tied to observable behavior. Those numbers are rolled up into tidy charts that suggest precision but often obscure the messy reality underneath.

The drift toward theater happens for understandable reasons. Maturity models offer structure and the appearance of objectivity. They let you say “we moved identity and access from level two to level three this year” and feel as if progress is measurable. Vendors and consultants reinforce the pattern by packaging assessments and dashboards around the model. Before long, the maturity score becomes an internal currency for status and budget, even though it seldom answers the questions that really haunt boards: which services will fail first under pressure, how quickly can you restore them, and what ugly trade-offs will you face in the middle of a serious incident.

The danger is not that maturity models are worthless. It is that they become self-referential. You end up optimizing for better scores, not better survivability. You congratulate yourself on a “green” domain even though the one team that actually runs the critical control in that domain is understaffed and one resignation away from crisis. You show overall improvement while the business quietly adds new dependencies that increase blast radius and recovery complexity. The model, which was meant to be a map, starts pretending to be the territory. When that happens, it is time to demote it.

Once you accept that “secure” is never finished, survivability becomes a more honest target. Survivability asks a different set of questions. When something important goes wrong, how does the organization bend? Which services degrade, and in what order? How quickly can you restore the functions that really matter to customers, regulators, and partners? This frame does not give up on prevention. It places prevention in the context of what happens when prevention fails, because at some point it always does.

The encouraging part is that survivability can be described in concrete terms at the level where boards operate. Instead of one synthetic cyber risk score, you look at demonstrated time to degrade a critical service under specific types of failure and demonstrated time to recover to a minimally acceptable state. You consider how long sensitive data might remain exposed or unreliable in a way that would break trust. You examine where your dependencies are concentrated: how many essential services hinge on a single identity platform, a single cloud region, or a single team. These are the kinds of measures that directors can debate intelligently because they live in the same space as other resilience topics, like physical outages or supply chain disruption.

Reorienting around survivability also changes your prioritization logic. Some projects that look modest on a maturity roadmap can have enormous impact on how the organization performs under stress. Shortening decision paths in an incident, simplifying a tangled dependency chain, or standing up clear degraded modes for a critical customer-facing function might not shift a level score, but they radically reduce the chance of catastrophic disruption. Conversely, some large technology deployments may produce impressive dashboard changes while leaving recovery behavior almost untouched. When survivability is your north star, the question “does this move our level?” is replaced by “does this make our worst realistic day meaningfully better?”

To aim at survivability, you need different data. Most organizations are very good at tracking what exists: policies, controls, assets, suppliers, and plans. Fewer track how those elements behave when the pressure is on. Measuring resilience means focusing on dynamic performance, not static presence. You stop asking only “do we have this control?” and start asking “what happened the last time this control was stressed or bypassed?”

Incidents are the first place to look for those answers. Every meaningful event reveals something about survivability. How long did it take to recognize that something was wrong? How quickly did the right people assemble? Where did handoffs stall? Which dependencies surprised everyone? How long were you operating in a degraded state, and what did that feel like for customers or internal users? Traditional post-incident reviews often focus on technical root cause and a list of remediation tasks. A resilience lens adds another layer of questions about time to degrade, time to recover, and the quality of communication under stress.

Exercises and experiments give you the other half of the picture. Tabletop scenarios, simulation drills, and carefully scoped chaos tests allow you to watch the organization practice failure in safer conditions. Instead of just confirming that a plan exists, you see whether people can execute it when information is incomplete and clocks are ticking. You notice where manual workarounds are brittle, where one key person dominates decision-making, or where tools slow responders down. Over time, you can turn those observations into repeatable measures: how long it takes to make a meaningful decision, how many critical services have tested degraded modes, how often rehearsals touch services that actually matter.

None of this will help if the board conversation stays stuck in the old script. Teaching directors to live with permanent incompleteness is as important as changing your metrics. That starts with reframing the core questions. Instead of walking into each meeting braced for “are we secure?”, you propose a small, consistent set of survivability questions you work through together: which critical services are currently most exposed to disruption, how well can you realistically recover them, and what losses are you consciously accepting in exchange for speed, innovation, or cost savings.

To support that kind of dialogue, you have to be intentional about what you show. One giant cyber index may feel efficient, but it hides the trade-offs that actually matter. A small set of resilience-oriented measures, presented alongside other enterprise risk indicators, does more real work. You highlight trends in demonstrated recovery performance, not just theoretical recovery time objectives. You call out single points of failure in terms that echo how the board already thinks about vendor concentration or logistics risk. You make the trade-offs explicit: shortening this recovery window will mean delaying that product launch or increasing spend in that domain.

Narrative matters here too. If every discussion of uncertainty feels like bad news, directors will keep asking you to “get to green.” You want to normalize the idea that cyber work is never complete while also showing that progress within that reality is possible and visible. That means telling stories where prior resilience investments paid off, even when the environment was hostile. You might talk about an incident where a critical service went into a planned degraded mode instead of failing outright, or where a recent exercise directly improved the speed and quality of a real response. Over time, the board learns to judge you less on whether incidents occur and more on how consistently you steer through them.

All of this depends on culture. You can define survivability metrics and update your slides, but if the organization still behaves as if incidents are rare embarrassments, the impact will be limited. A resilience-focused culture assumes that hits are coming and treats preparation as part of normal work. It sets the expectation that surprises are data, not shame. It makes room in schedules and roadmaps for drills, simplification, and reflection, rather than stuffing every spare hour with new projects.

In practice, that culture looks like regular cross-functional exercises that pull in people from across technology, operations, and the business, not just a small incident response core. It looks like post-incident reviews that focus on how systems and processes set people up, rather than who made the last mistake. It looks like leadership carving out time and budget for work that reduces fragility, even when it does not come with a shiny new tool or a quick maturity bump. The way you use metrics reinforces this. You celebrate improvements in recovery performance and decision speed as achievements, not just clean audits.

Incentives are the lever that make these patterns stick. If career success depends on spotless dashboards and never admitting uncertainty, people will hide weak spots until they explode into the open. If success includes honest reporting of vulnerabilities, active participation in exercises, and clear documentation of risk trade-offs, the organization becomes more transparent about where it can take a hit and where it is exposed. Over time, the combination of practice, transparency, and aligned incentives builds a quieter, deeper confidence than any single metric. Teams know they will be tested. Boards know the program is tuned for survival, not just optics.

For a Chief Information Security Officer (C I S O) or technology executive, this shift from perfection to resilience is a leadership choice as much as a technical one. It means giving up the comfort of promising that things will someday be “done” and replacing it with a more mature promise: that you will continually improve the organization’s ability to withstand and recover from real-world events, and that you will be honest about where it still falls short. It means inviting your peers and your board into a more grown-up risk conversation where incompleteness is expected, where trade-offs are explicit, and where preparation for impact is as valued as prevention.

A simple way to start is with one question. Ask your leadership team, and then your directors, to walk through what they honestly believe would happen if your two most critical services took a serious hit in the next quarter. How quickly would you know? Who would be in the room? What would customers experience? How long would it take to get back to an acceptable state? The gaps, disagreements, and surprises in that discussion are not a failure. They are your roadmap.

When you move the story from “Are we secure?” to “How well can we survive the hits we know are coming?”, you do not lower the bar. You move it to where the real game is played. That is where resilience lives. And that is where leaders can actually lead.

Resilience over Perfection
Broadcast by