Newsletter Article

The Security Data Cost Problem Is Really an Architecture Problem

For years, the response to improving security visibility was relatively straightforward: add another log source. Then another. Then another.

The modern enterprise now generates security telemetry from endpoints, identity systems, SaaS applications, cloud infrastructure, firewalls, APIs, containers, development platforms, business applications, AI services and network infrastructure. Those environments continue to grow.

The problem is not that those logs lack value.

Snare Insider Newsletter Series Article

Make sure you Subscribe

The problem is assuming every log needs to be processed, retained and queried in exactly the same place.

Many security platforms price some combination of ingestion, storage, processing, search or workload consumption. Ingestion-based pricing remains common, although structures vary considerably between vendors. At enterprise scale those economics matter: a security organisation ingesting tens or hundreds of terabytes each day experiences the impact of a small unit-price change very differently from a smaller SOC.

This creates the SIEM cost paradox we explored in Issue 12:

Better visibility requires more security evidence, but indiscriminately sending more evidence into the highest-cost tier can make visibility financially unsustainable.

There is a structural reason ingestion pricing bites so hard, and it is worth naming precisely. Ingestion pricing couples two unrelated things: the value of a log at the moment it arrives, and the cost of keeping it available months later. Most security telemetry has a steep, short value curve for detection and a flat, long value curve for investigation. Pricing both through the same tier forces a decision that is necessarily wrong for one of them.

The answer is not simply “collect less.” That creates investigation gaps, and, as Berlin and Thomson Reuters both illustrate, the gap does not become visible until the moment it is most expensive.

Where does each class of data need to go?

Data class Examples Detection Investigation Route it to Working retention
High-signal alerts EDR detections, IdP risk events, DLP alerts High High SIEM, real time SIEM hot tier
Authentication & privilege IdP sign-ins, 4624 / 4625 / 4672 / 4720 / 4728, PAM session records High Very high SIEM plus independent retention 12 months+ searchable
Endpoint process & command line 4688 with command line, Sysmon 1 Medium Very high Filtered subset to SIEM; full fidelity to independent retention 90 days+ central, extendable on demand
Egress & file access Proxy and firewall egress, 5145, NetFlow / IPFIX Medium Very high, drives scope and notification Summarised to SIEM; full record to retention 12 months+
Administration platform audit RMM, hypervisor, backup console, CI/CD, PAM Medium Very high Independent retention first; alert subset to SIEM 12 months+
High-volume, low-signal DNS, verbose web access, debug and diagnostic Low per event High in aggregate Retention and lake; sampled or aggregated to SIEM 90 days – 12 months
Duplicate or redundant The same event arriving via two collectors None None De-duplicate upstream, before billing n/a

Filtering, transformation and routing are not the same thing

Most SIEM cost-reduction programmes stall because they treat three distinct operations as one. Separating them is what makes cost reduction compatible with better investigation coverage rather than opposed to it.

FILTER Decide what a given destination receives. The event still exists; one consumer simply does not get it.
TRANSFORM Reshape or reduce an event, drop verbose fields, normalise to a common schema, aggregate repeated events, without losing the record itself.
ROUTE Send the same event to more than one destination, with different filtering, transformation and retention applied to each.

A cost programme that only does the first is a visibility-reduction programme wearing a finance label.

Retention has two clocks

Retention decisions are usually framed against a single number, a compliance minimum, or whatever the SIEM licence affords. In practice there are two independent clocks, and they pull in different directions.

The first is the investigation clock, set by dwell time and detection lag: 247 days on IBM’s current average; March to 30 June in the Thomson Reuters disclosure. This clock determines how far back the evidence must reach.

The second is the regulatory clock, which runs on a much shorter fuse but demands answers about a much earlier period.

Obligation Clock What it implies for logging
GDPR Article 33 Notification to the supervisory authority without undue delay, and where feasible within 72 hours of becoming aware Scope of personal data affected must be established in days, from evidence that may be months old
NIS2 Article 23 Early warning within 24 hours; incident notification within 72 hours; final report within one month The one-month final report requires a reconstructed root-cause timeline, not just an impact statement
DORA Article 19, per Commission Delegated Regulation (EU) 2025/301 Art. 5 Initial notification within 4 hours of classifying an incident as major, and no later than 24 hours from awareness; intermediate report within 72 hours; final report within one month The 4-hour clock starts at classification, which requires enough evidence to classify, immediately
SEC cybersecurity disclosure, Form 8-K Item 1.05 Within four business days of determining an incident is material Materiality determination depends on knowing scope; scope depends on retained egress and access evidence
ISO/IEC 27001:2022 Annex A 8.15 and 8.16 Logging, and monitoring activities, as stated controls Auditors increasingly test whether logs are protected and retrievable, not merely that logging is enabled
ASD ISM, ISM-1988 (replacing the rescinded ISM-0859) Event logs retained in a searchable manner for at least 12 months, with records-disposal alignment to NAA AFDA Express v2 The December 2024 change removed the seven-year figure and added the word “searchable”, the requirement is now about retrievability, not storage
ASD ISM, ISM-1815 Event logs protected from unauthorised modification and deletion (Essential Eight ML2 and ML3) Logs must survive compromise of the host that produced them

The ISM change is worth dwelling on, because it captures the shift precisely. ASD rescinded ISM-0859, the long-standing control requiring most event logs to be retained for at least seven years, in December 2024, and replaced it with a requirement that logs be retained in a searchable manner for at least 12 months.

The regulator moved the emphasis from how long you keep it to whether you can actually find it. That is an architecture requirement, not a storage requirement.

None of these clocks care where your data physically lives. They care whether you can answer, within hours, a question about something that happened months ago. “Retained” and “retrievable inside the notification window” are different architectural properties, and only one of them is usually specified in a SIEM contract.

The collection and routing layer should be able to make those decisions before the data reaches the most expensive destination, and before the clock starts.

Snare Solutions
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.