Key Takeaways
|
We just finished the first run of our AI-Augmented SOC roadshow. Nashville, Atlanta, Toronto. Two live demos in every room, one on agentic triage in Google SecOps and one on cloud remediation in Wiz, with a fifteen minute segment from me in the middle on log source rationalization.
The conversation that followed me out of every room was about something interesting, and it arrived as the same question from CISOs, platform owners and analysts: what does this need from us before it works?
That question deserves a better answer than this market usually gives it. Here’s ours.
The number that starts every one of these conversations
We model a mid-size enterprise estate at 916 gigabytes of security telemetry per day across 48 distinct sources. At a modeled rate of $1,150 per sustained gigabyte per day per year, covering ingest and retention, that is roughly $1.05 million a year before anyone has written a detection. That rate is an input, not a claim about your contract. Any account large enough to be having this conversation is on a negotiated rate card, so in the workshop we replace it with yours and every figure below moves with it.
The interesting number is the next one.
Three sources account for 700 gigabytes of that total, which is seventy-six percent of the bill. Sysmon at 390 gigabytes, next-generation firewall traffic logs at 163, and the Windows Security event log on servers at 146.
Almost every security leader I have shown that to reaches for the same answer, which is to cut the three big ones. It isn’t, and the reason why is where this gets interesting.
Coverage is a union, and that changes the shape of everything
This is the technical heart of it, and the part most cost-reduction exercises get wrong.
We score every log source against MITRE ATT&CK Enterprise v19.2: 697 techniques and sub-techniques, 222 and 475 respectively. The technique-to-source mapping is MITRE's own, taken from the log source references published in the framework rather than from our assertion about what a product can see.
Of those 697, 652 are reachable by some log source. The other 45 have no telemetry mapped to them at all, which is an honest finding rather than a gap in anyone's product.
Now the important bit. Coverage is a union, not a sum. A technique observed by four different sources counts once. You do not get four units of coverage, you get one technique you can see, plus three sources of corroboration and three ingest bills.
That single property is why cost and coverage have completely different shapes as you add sources. Cost climbs in a straight line, because every gigabyte bills the same as the last one. Coverage climbs steeply and then flattens into a wall, because the fifth source watching process creation adds almost nothing the first four did not already give you.
The practical consequence is measurable. Take the twelve most efficient sources in the catalog, ranked by techniques observed per gigabyte per day. Together they are 17 gigabytes, under two percent of the estate, roughly $20,000 a year. They cover 522 of the 652 reachable techniques, eighty percent of the total.
Eighty percent of the coverage for two percent of the cost. That is what the union does to the shape of the curve, and you can reproduce it against any estate.
One precision point, because this is where the argument gets misquoted. Coverage here means a technique has at least one source in the estate capable of observing it. It does not mean you have a tuned, tested detection firing on it. Observability is the ceiling; detection content is the work you do underneath that ceiling. The reason to measure the ceiling first is that no amount of detection engineering gets you above it.
Yield per gigabyte, and the single best example in security
Once you accept that coverage is a union, the right unit of measure stops being volume and becomes techniques observed per gigabyte per day. Run the whole catalog through that lens and the spread is larger than almost anyone expects.
At the top of the catalog:
• External attack surface management: 553 techniques per gigabyte
• Network device configuration and CLI logging: 273
• Entra ID audit logs: 137
At the bottom:
• Next-generation firewall traffic logs: 0.33
• Full packet capture from network detection and response: 0.23
More than three orders of magnitude between the best and worst sources inside the same estate, and almost nobody is measuring it.
The obvious objection is that a small denominator flatters the ratio, and it does. External attack surface management scores 553 because it is fifty megabytes a day, so any mapping at all produces a large number. That is the point rather than a flaw in the metric: a source contributing 27 techniques for fifty megabytes is the easiest decision in the catalog. Yield per gigabyte is the right measure when volume is the constraint, and the wrong measure for deciding whether a detection is any good. We do not use it for that.
But the example I keep coming back to, the one that landed hardest in all three rooms, involves no comparison between products at all. It is the same log file in two places.
• Windows Security event log, domain controller: 7.8 GB per day, 280 techniques, 35.8 techniques per gigabyte
• Windows Security event log, server: 146.5 GB per day, the same 280 techniques, 1.9 techniques per gigabyte
ATT&CK maps techniques to the log type rather than to the host, so both rows carry the same 280. That is exactly why the comparison is useful: the mapping is identical, the parser is identical, and the only variable left is where you collected it.
Same log, same event IDs, same parser. Nineteen times the efficiency, decided entirely by which host tier you collect it from.
No product decision appears in that comparison and there is no vendor to blame. It is a collection design decision, and most organizations made it years ago by accident, by turning on a collector policy at the domain level and never revisiting it. Event 4662 carrying the DS-Replication-Get-Changes GUID only appears on a domain controller. A replication request from anything that is not a domain controller is DCSync, and only the DC is positioned to record it. Collecting that same log from nine hundred member servers buys you the bill rather than nine hundred times the signal.
Why Log Source Rationalization Matters for an Agentic SOC
This is where a cost conversation turns into a capability conversation, and it is why we put the segment in front of the demos rather than after them.
An agent cannot investigate evidence that was never collected, corroborate a finding only one source saw, or reason about a technique your estate has no telemetry for. That is not a limitation of the agent. It is arithmetic about what sits in the platform when the agent goes looking.
Two numbers from our model make it concrete.
First, 45 of the 697 techniques have no log source mapped to them in ATT&CK at all. No product covers those today, ours or anyone's. Any vendor who tells you their agent has full coverage of the framework either has not read the framework or is counting something other than what you think.
Second, and more actionable: in a well-built detection tier of 34 sources covering 617 techniques, 126 of those techniques rest on a single source. One collector fails, one licensing change lands, one team decommissions a server, and a fifth of your detection coverage goes dark silently. An agent running on that estate will confidently triage the alerts it receives and will never tell you about the alerts that stopped arriving.
That fragility is invisible in every dashboard I have been shown. It appears only when you compute coverage as a union and then ask how many techniques have exactly one witness.
Which means your log source decisions are no longer a finance conversation. They are the ceiling on what your AI can do.
Where ATT&CK is the wrong yardstick, and why that matters for Wiz
Seven sources in our catalog score zero on technique coverage, because ATT&CK maps no log source to any of them:
Wiz issues and detections, vulnerability scanner findings, threat intelligence platform indicator matches, backup and recovery telemetry, secrets management and vault audit, physical access control and badge events, and deception and canary tokens.
Together they are 0.88 gigabytes per day, under a tenth of one percent of the estate.
A naive reading of a coverage model says those are the first things to cut. That reading is wrong, and the error is in the yardstick rather than in the sources.
ATT&CK describes adversary behavior. It scores a source by whether that source witnesses an attacker doing something. Wiz answers the question an alert cannot: was that workload reachable from the internet, was it vulnerable, and does it hold anything worth taking. An alert on a container says something happened. Wiz decides whether it mattered.
That is a prioritization contribution rather than a coverage one, and it gets measured in analyst hours not spent and incidents not escalated. Same for a vulnerability scanner. Same for backup telemetry, where a job failure or a deletion of snapshots is a known precursor to ransomware deployment and arrives before encryption does. Same for a canary token, which produces almost no volume and almost no false positives, and which most organizations never deploy because it is not a product category anybody budgets for.
So we tier those sources on a different basis, and we say so out loud. Scoring them at zero and keeping them anyway is the model being honest about the one thing it measures.
Four tiers, one question each
The framework we use is deliberately small, because a taxonomy nobody can recite is a taxonomy nobody applies. Four tiers. Ask the questions in order and stop at the first yes.
| TIER | THE QUESTION | WHERE IT LIVES |
| Detection | Would losing this blind a detection you run today? | Hot, parsed, searchable, retained for the detection window |
| Investigation |
Would you want it the hour an incident opens? |
Warm and queryable. Real value once something is happening, wrong economics for hot storage |
| Compliance |
Does a regulation or an auditor name it? |
Cold archive. It should never touch your SIEM ingest bill |
| Drop or sample |
Is the honest answer to all three no? |
Off, or sampled to a fraction |
One clarification that matters, because it is where the framework gets misread. Detection tier does not mean untouched. Sysmon is 390 gigabytes and 423 techniques in our model, more than any other single source. It earns the detection tier comfortably, and it is still the first thing we tune. The verb for your largest detection sources is configure, not delete. A tuned configuration that drops the noisiest, lowest-value event classes takes real volume out without touching the detections that depend on it. We do not publish a percentage for that, because it depends on what your config is already collecting, and anyone who quotes you a number before reading it is guessing.
What the honest verdict looks like
Run the full 48-source estate through that model and here is what comes out.
| TIER | SOURCES | VOLUME | SHARE OF ESTATE |
| Detection | 34 | 617 GB/day | 67.4% |
| Investigation | 8 | 13 GB/day | 1.4% |
| Compliance | 2 | 25 GB/day | 2.8% |
| Drop and sample | 4 | 261 GB/day | 28.5% |
Now cost each tier at the rate its storage class deserves rather than at the hot rate. In this model detection bills at full rate, investigation at roughly a third, compliance at cold-archive rates near six percent, and the drop tier at zero. The estate goes from 916 gigabytes at hot rates to 623 gigabytes costed, which is $1.05 million down to about $716,000. A 32 percent reduction.
The price of that reduction is ten techniques, out of 633 the catalog covers.
And here is the part I like best, because it is the part that’s checkable. Those ten techniques are not scattered across the kill chain. All ten sit in Reconnaissance and Resource Development. Active Scanning, Scanning IP Blocks, Vulnerability Scanning, Wordlist Scanning, Gather Victim Identity Information, Email Addresses, Establish Accounts, Compromise Accounts, and two social media sub-techniques.
Every one of them is pre-attack. Nothing moves in Execution, Credential Access, Lateral Movement or Impact. You are giving up visibility into adversaries researching you on the open internet, which is real but is the cheapest thing on the list to lose and the hardest thing to act on anyway.
That is the trade, stated plainly, with the cost named. A third off the bill for ten pre-attack techniques.
What this means for the stack
First the honest caveat. Everything above is platform-agnostic by construction. The union property, the yield metric, the four tiers and the fragility count are arithmetic over your own estate, and none of them depends on which SIEM the data lands in. Arctiq is not a single-platform shop and the rationalization work stands on its own against whatever you run today. What follows is the stack this roadshow was built around and the one we would argue for right now, not a claim that the method requires it.
Google SecOps is the data plane and the detection engine. It is where the tiering decisions become real, where the detection catalog lives, and where the agents run. It is also where the economics of this conversation get decided, because the ingest tier is the thing being rationalized.
Wiz is the context layer ATT&CK cannot score and every alert queue needs: exposure, reachability, and what the asset is worth. That is the difference between an alert and a prioritized alert, and it earns more per gigabyte than anything else in the estate precisely because it is not measured in gigabytes.
Google Threat Intelligence turns a detection into an attribution and a priority, which is what lets a triage decision happen in sixty seconds rather than get deferred to a human who will reach it tomorrow.
Those three together are a credible agentic SOC. None of them substitutes for knowing what is in your estate and why. The agents are the most sensitive of the three to that groundwork, because an agent's reasoning is bounded by the completeness of the evidence in front of it in a way a human analyst's is not. A human knows when a log is missing. An agent sees an absence of evidence and reports an absence of findings.
The ask
If any of this describes your estate, we run a half-day Log Source Rationalization Workshop and we run it on your numbers rather than ours.
Real estates are messier than the model. Ours is 48 clean sources; yours probably has two SIEMs mid-migration and a dozen collectors nobody has looked at since the person who built them left. Half a day gets you through the sources you can name. The ones you cannot name are their own finding, and usually the most interesting one in the room.
You bring your source inventory and your last invoice. We bring the model, the 697-technique coverage map, and the tiering method. By the end of the session you have four things, scoped to the sources your inventory actually names:
- A tiered catalog of every source you ingest, with the reason for each tier decision written down
- Your own cost model, which you can change in front of your CFO and watch the number move
- A defensible drop list with the coverage cost of each item named rather than estimated
- The fragility list: which techniques in that inventory rest on a single source today
That output feeds into a SOC Optimization and Visibility Workshop, where the tiered estate becomes a detection catalog mapped to ATT&CK, alert volume gets rebuilt around what an analyst can close in a shift, and the agents get pointed at an alert population worth automating.
The order matters. Applied to clean, tiered, well-understood telemetry, agentic AI is leverage in the real sense of the word. Applied to an unrationalized estate, it is an expensive way to automate noise faster.
Bring us your inventory and your invoice. We will build you the same picture you just read, on your own numbers, before you spend another quarter paying for telemetry that has never produced a detection. Connect with us to see how Arctiq can help you build a stronger foundation for your agentic SOC.
Tags:
Cybersecurity
September 30, 2026