Zero Signal
The renewal meeting starts the way it always does. Someone pulls up a deck showing last year’s spend on observability and security analytics, then the new “optimized” proposal for the coming term. A finance leader points to the rising number and asks why the cost keeps climbing when the number of material incidents has not gone down. Security leaders insist the logs are essential. Vendors promise “unparalleled visibility.” Yet when someone asks which specific risks would increase if they dialed anything back, the room goes quiet. This narration is part of the Wednesday “Headline” feature from Bare Metal Cyber Magazine, developed by Bare Metal Cyber, and it lives in that quiet gap between spend and specificity.
Over the last decade, security telemetry has turned from a side effect into a structural cost of running digital infrastructure. Logs, metrics, traces, endpoint events, and identity signals are all streaming out of your systems at high volume. It is no longer just a matter of paying for a few extra servers or a slightly larger storage bucket. Entire pricing models for cloud services and Security information and event management (S I E M) platforms are built around the amount of data you ingest and retain. The reflex has been simple: log everything, keep it for a long time, and hope that any important signal is somewhere inside that mass of data. For many organizations, the economics and the human attention required to sustain that pattern now outweigh the risk it actually mitigates.
When leaders look ahead a year or two, uncomfortable questions are waiting. Which events do they really need to see in near real time to catch the threats that matter most. Which data do they only need in a reconstruction after the fact. Where is telemetry truly a regulatory requirement, and where is it more of a comfort blanket that nobody wants to touch. This story walks through those questions as a design problem rather than a moral one. The goal is not to shame teams for collecting too much or too little, but to make telemetry a deliberate instrument in the way you manage risk, cost, and human focus.
The situation most organizations are in did not start with a grand plan. It started with small, reasonable choices. A development team once turned on extra logging to debug a nasty production issue and never turned it back down. A cloud platform shipped with generous “recommended” audit trails that were left at defaults. A new endpoint product promised richer events “at no extra effort,” so someone clicked the checkbox. Over time, those choices compounded. What began as a few targeted feeds grew into a mesh of signals from every layer of your stack, all flowing into one or more analytics platforms. Each addition was rational on its own, but nobody ever stepped back to ask what the whole stream looked like or what it was for.
The day to day experience of this growth is familiar. Dashboards are full of activity, but genuinely new or actionable insights are rare. Analysts context switch between several tools to answer basic questions because no single view was ever designed to support the way they investigate. Engineers and site reliability staff waste time hunting through long lists of fields, trying to remember which variation of an event actually has the data they need. Over time, noise stops being a technical issue and becomes a cultural one. People quietly stop believing that an alert is meaningful. They stop expecting a dashboard to tell them something they did not already know. Decisions drift back toward instinct, anecdote, and the loudest voice in the room, because the telemetry is no longer a shared decision instrument.
Several invisible forces hold this pattern in place. Compliance language that says “retain logs for a period of time” gets simplified into “retain everything for that period, just to be safe.” After each incident, retrospective conversations end with a familiar promise to add more feeds so that “we never miss this again,” but rarely remove anything else to compensate. Vendors design their business models to reward ingestion growth, then pitch new analytics and assistants that assume you will keep every possible event. Over time, “collect by default” stops being a choice and becomes an unwritten norm. The organization loses the habit of asking a basic question: what do we expect to learn from this data, and how would we notice if it quietly stopped flowing.
Once leaders acknowledge how much they are spending on telemetry, the first instinct is often to look for quick cuts. That reaction is understandable, but it hides a different kind of risk. Not all data is equal. Some logs are harmless exhaust that can be summarized or dropped entirely. Others are the only way you will ever reconstruct how an intruder moved through a privilege path or tampered with a key business process. The hard truth is that most environments were never instrumented outward from a clear view of risk. They grew from defaults, vendor templates, and local debugging needs. That means there is rarely a shared map of which signals are tied to the risks the organization truly cannot afford to miss.
You can see the consequences in very different kinds of organizations. A cloud-native product company might retain verbose logging for multiple test and development environments, because it was easy to turn on when engineers were troubleshooting. At the same time, the production customer portal might only have simple access logs and no deeper view into what users did once they authenticated. Telemetry spend is dominated by noise from non-critical workloads, while the trail that would reveal account takeover patterns is thin. A highly regulated financial institution might keep years of low-value virtual private network logs out of habit, while visibility into privileged actions on core payment systems is fragmentary and hard to correlate. In both cases, leaders are paying for volume in the wrong places while carrying blind spots where failures would be existential.
The key shift is to treat telemetry as a risk instrument that you shape to the threats and obligations you actually own. That begins with a simple inventory of what really matters. Identify the business services that generate revenue or carry regulatory obligations. Map the data stores that contain sensitive information and the identity paths that grant or broker powerful access. For each of those, ask a few direct questions. If something went wrong here, what would we need to see within minutes. What would we need to reconstruct within hours or days. What evidence would we owe our customers, our regulators, or our own board. Those answers point to specific events and specific levels of detail that matter, instead of an undifferentiated call to “collect more.”
When you frame the problem that way, it becomes clear that different parts of your environment deserve different levels of telemetry. One practical pattern is to define a top tier of systems and services that are genuinely business critical. These might be customer-facing portals, core transaction engines, or identity platforms that everything else depends on. For this tier, you invest in deliberate, often custom, instrumentation that is tied to threat models and to real investigation workflows. You decide which fields must appear in each event, how long you need detailed records to be searchable, and what “good enough” looks like for correlation across tools.
Below that top tier, you can define other levels that rely more on summaries, platform defaults, or shorter retention. Internal administrative tools with low risk might only need simple access logs kept for a modest period. Ephemeral test workloads might only emit health metrics and error counts, with no detailed traces stored anywhere central. Batch processing jobs might be adequately covered by one completion log and an aggregate report. The important point is that each level reflects an explicit decision about risk and response, not a blanket acceptance of whatever a vendor suggests as a default.
Designing telemetry on purpose also requires clear ownership and simple governance. Someone needs to be responsible for the overall telemetry portfolio, not just each individual tool. That responsibility includes deciding who can add new sources, who approves changes to retention, and how conflicts between cost and coverage are resolved. It also includes a basic design review discipline. When a new service is built or an existing one is significantly changed, part of the review should cover what events matter for detection, where those events will land, and who is accountable for watching their health over time. These questions do not have to be formal or bureaucratic, but they should be asked consistently.
A useful artifact for critical services is a short telemetry design note that can be read in a few minutes. It states the main risks the service carries, the key questions investigators need to answer during an incident, and the specific signals that support those answers. It also notes where the data is stored, how long it is kept, and which teams are responsible for keeping it usable. Such a note is not a policy museum piece. It is a living reference that can be updated when threats change or when the cost of a particular signal no longer matches its value. Over time, a small library of these notes gives leaders a much clearer picture of what they are paying for and why.
Culture is the thread that runs through all of this. “Log everything” is rarely written down as a policy, but it is often present as a defensive mindset. Engineers worry that if they turn off a verbose log and something goes wrong later, the blame will land on them. Security teams fear being told that they “missed” an attack because they did not insist on collecting every possible field. Vendors and consultants sometimes encourage this fear, because the safest path for them is to say that more data is always better. Changing this culture does not mean encouraging recklessness. It means rewarding clear reasoning about risk instead of sheer volume.
Leaders can help by changing the kinds of questions they ask in reviews and incident follow-ups. Instead of asking why a particular log was not enabled, they can ask what question the team needed to answer and which options they considered to support that. Instead of praising teams for adding yet another feed, they can praise teams for retiring a noisy source that no longer justifies its cost. Over time, this changes the incentive structure. Teams learn that they will be backed when they make thoughtful trade-offs, not just when they protect themselves by over-collecting.
Budget conversations also change when leaders see telemetry as one design lever among many. When the cost line climbs, the question is no longer “why is our S I E M bill so high” in isolation. It becomes “what share of our security investment is tied up in collecting and analyzing data, and is that balance still right.” Leaders can compare the spend on telemetry against investments in identity controls, architecture simplification, or incident response training. In some cases, trimming low-value logs and shortening retention in non-critical areas can free enough budget to fund a long-delayed improvement elsewhere that has more direct impact on risk.
The same is true for human capacity. Every new source of telemetry comes with work attached. Fields must be normalized. Parsers must be built and maintained. Detections must be tuned. Dashboards must be owned. False positives must be understood and refined. When leaders ask for an honest accounting of how much analyst and engineering time is being spent on these tasks, they often find that a large fraction of their attention is locked up in keeping low-value signals alive. That attention could instead support deeper investigation skill, closer collaboration with product teams, or proactive threat hunting in the parts of the environment that matter most.
When you pull all of these threads together, the core idea is simple. Telemetry is not a sacred obligation to collect everything forever. It is a tool for answering specific questions about systems, people, and data under stress. At its heart, this story is about treating that tool as something you design and adjust, rather than something that just happens to you. The leadership meeting that opened our journey looks very different when people can say with confidence which services live in the top tier, which events support critical detections, and which parts of the current bill are noise that everyone is willing to release.
For leaders who internalize this view, a few behaviors change. They stop tolerating unchecked growth in telemetry just because it feels safer. They expect teams to explain which risks a new data source addresses and how they will know when it stops adding value. They ask vendors to justify not just capabilities but also the economic trade-offs of keeping high-volume feeds online. They use their influence with auditors and boards to describe a risk-based strategy for visibility, rather than defaulting to simple volume metrics.
Most importantly, they bring better questions into their own discussions. Which business journeys truly deserve tier-one telemetry. What evidence would we need, and from where, if a serious incident unfolded in that space. Where are we currently paying for detail that we do not use, and how could that capacity be re-deployed to areas that are clearly under-served. When those questions become routine, they do more to strengthen real visibility than any blanket instruction to “log everything” ever could.