Most threat hunting programs produce activity theater

MKMarta K. · Senior Detection Engineer & Incident Responder
Threat Hunting·10 min read

The threat hunting program I inherited looked healthy on a slide: a hunt calendar, ATT&CK coverage in the high seventies, tidy quarterly reports. Not one hunt had changed a detection in eighteen months. This is how I separate a real hunt from a rebranded IOC sweep.

The threat hunting program I inherited two years ago had a Confluence space and a quarterly hunt calendar. A slide for the chief information security officer (CISO) showed ATT&CK coverage in the high seventies. No detection had changed as a result of a hunt in the previous eighteen months.

Every "hunt" in the log followed the same script: an analyst took the indicator of compromise (IOC) list from a vendor's threat report, ran the hashes and Internet Protocol (IP) addresses through the security information and event management (SIEM) platform, got zero hits, and closed the ticket. From the CISO's seat it looked like hunting. From mine it looked like a cron job with a byline.

That pattern is common. Only 15.6% of organizations self-assessed as very mature and hypothesis-based in the SANS 2022 survey. Claiming maturity while running sweeps is the industry's default state.

In Brief:

  • Scheduled IOC sweeps, ATT&CK coverage percentages, and vendor alert triage are the three most common performances of threat hunting. All three produce activity theater without testing a hypothesis.
  • A real hunt starts from a hypothesis about specific adversary behavior and accepts a null result. It ships a changed detection or closes a visibility gap.
  • Programs drift into performance because activity is countable and outcomes resist counting. Hunt counts survive quarterly reviews. Hypothesis discipline is harder to keep.
  • Doing it requires protected time off the alert queue, a named owner for every hypothesis, and a formal handoff into the detection engineering backlog.

What a performing hunting program looks like

A performing program produces recurring artifacts that photograph well in a quarterly review and change nothing about what the security operations center (SOC) can detect.

Scheduled IOC sweeps rebranded as hunts

The PEAK framework defines threat hunting as any manual or machine-assisted process meant to find incidents that automated detection missed. That breadth is deliberate, and PEAK recognizes baseline and model-assisted hunts alongside hypothesis-driven ones.

Its co-author David Bianco has pushed back on the emphasis anyway: finding incidents is not the real point, because automated systems already do it better than analysts can.

For running a program rather than defining the field, I use a stricter test. A hunt is a hypothesis about adversary behavior, tested against data, that produces an improvement to automated detection whether or not it finds a threat, and everything here measures against that narrower standard.

The TaHiTI methodology, built by the Dutch financial sector, is explicit that threat hunting is not searching for indicators of compromise and not simply running a query in a tool: if a tool can do the work autonomously, it is not hunting.

Bianco's Hunting Maturity Model puts indicator searches through historical data at HMM1, the minimal rung one step above pure alert resolution. Known indicators belong inside automated detection and response, firing continuously rather than once a quarter. An IOC sweep tests no hypothesis and requires no knowledge of the environment, so when it comes up empty, it teaches nothing.

MITRE coverage percentages standing in for findings

I have written before that a green ATT&CK heatmap counts how many rules you have written without testing whether any of them work. Researchers presenting at USENIX Security 2024 found that endpoint detection and response (EDR) vendors hover at 48% to 55% coverage, and that many mapped techniques exist only as low-priority rules the vendors themselves define as unreliable.

The same research found customers read that coverage as a definite measure of security, treating 90% coverage of MITRE ATT&CK as 90% secure. Binary coverage mapping collapses the difference between a tuned detection that fires on real activity and a noisy rule that got tagged with a technique ID once.

MITRE tells people not to use its framework this way. Its own guidance warns against chasing 100% coverage and against declaring victory on the strength of a single mapped technique. A coverage percentage is an inventory of rules. When it stands in for hunt findings on a slide, hunting has been replaced by accounting.

Vendor hunts that are really alert triage

The UK Home Office guide flags the tell in vendor service descriptions: heavy references to known threats and common attack techniques imply a reactive stance better suited to protective monitoring than to proactive searching for the unknown. That reactive stance is the industry baseline.

Read your provider's monthly hunt report the same way. If almost every line item traces back to an alert their platform raised, you are mostly seeing triage, not proactive hunting. Expel, for example, places IOC sweeps within essential security hygiene and describes them as one component of effective threat hunting. A report whose line items all stop at triage is documenting duplicated monitoring, not hunting.

What separates performance from practice

The dividing line is not tooling or headcount but whether the work starts from a falsifiable claim, tolerates a null result, and ends by changing a detection.

A hunt starts from a hypothesis about a specific adversary behavior

TaHiTI encodes the idea in its name: targeted hunting, not random browsing through logs. A usable hypothesis names the behavior and the place, and it explains the reason. "An attacker holding valid credentials is using built-in admin tooling to move through our cloud estate without tripping any current alert" is a hypothesis. "Check the new IOC feed" is a chore.

The scoping test comes from incident responder Lesley Carhart, who argues that every hypothesis should be formed so that you can either confidently falsify it or prove it is happening and initiate incident response. That means the analyst defines in advance what evidence would disprove the hypothesis. If a hunt cannot fail, it is not one.

Real hunting is willing to find nothing

TaHiTI treats both proven and disproven hypotheses as valid outcomes, and reserves a separate "failed" verdict for hunts where the data needed to test the claim did not exist. A disproof is a success.

Bianco is even more blunt about the incident-count trap: measuring a hunting program by the number of incidents it opens is a terrible idea, because hunters do not control the threat actors.

A recent public example is a July 2025 hunt that the Cybersecurity and Infrastructure Security Agency (CISA) and the US Coast Guard ran at a critical infrastructure organization. It searched host, network, and cloud data, mapped findings to 18 MITRE ATT&CK techniques, and found no malicious actor. It still identified six significant security gaps.

A program that cannot report a null result honestly will manufacture positive-looking activity instead, which is how sweeps and heatmaps take over.

The output changes a detection

PEAK assigns the hunt's organizational value to its Act phase, where findings become production detection rules so the team stops hunting the same behavior over and over. That output is measurable without gaming.

The UK Home Office guide lists the percentage of successful hunts that result in a new detection analytic or rule as an explicit service-level metric. A useful hunt report ends with a rule or a closed gap. A documented disproof also delivers value to the SOC.

Why programs drift into performance

Programs drift into performance through measurement.

Performance is measurable and hunting resists it, so measurement wins

The FIRST.org metrics catalog names the failure: Goodhart's Law is the most common way metrics programs quietly degrade, as the number becomes the goal rather than the outcome it was meant to represent. Hunt counts and sweep completions can be tallied by anyone, weekly, with no judgment required.

Coverage percentages can too. "Detections created from hunts" requires hypothesis discipline and a working handoff to detection engineering, so under quarterly reporting pressure the countable thing wins. Independent consultant Kostas T. calls the per-hunter hunt count a bad metric that rewards quantity over quality and pushes hunters to rush.

The drift compounds over time. Hand-tracked effectiveness does not survive a reorg. When the senior hunter leaves, the sweep schedule keeps running, and quietly becomes the program.

What doing it actually requires

Protect hunting time and assign each hypothesis an owner. Then create a formal path into detection engineering.

Time carved out of the alert queue, and an owner for the hypothesis

Effective threat hunting programs assign analysts' time explicitly for hunting. Target learned the rotation trap publicly: before its program refresh it had no full-time hunters and pulled SOC analysts into week-long hunts on an eight-week rotation, which the company said often led to operational inefficiencies.

When I rebuilt the program I inherited, I used the existing tooling and protected calendar time. Two analysts came fully off the alert rotation on Tuesdays and Wednesdays, no exceptions short of a declared incident, because a hunter who can be paged back to the queue is a triage analyst with a side project.

The second change was a name on every hypothesis before the hunt block started. The owner writes the hypothesis and defines what would falsify it. The owner also signs the outcome, whether that is an escalation, a disproof, or a data gap. Ownerless hunts default to whichever query is easiest to run, which is how IOC sweeps sneak back into the program.

A feedback loop from hunt to detection engineering

SpecterOps describes the detection engineering backlog as an input chokepoint for the whole detection and response program, with threat hunting as a formal channel feeding it hypotheses, research, and queries.

Document the handoff as a formal artifact. Detection logic, required telemetry, expected behavior, and false positive considerations should be written down well enough that an engineer who was not on the hunt can build and tune the rule.

Then measure the program on what crossed that boundary. A solid basket covers detections created or updated, incidents opened, gaps identified and closed, misconfigurations found, and techniques hunted. The program I inherited now reports detections shipped and gaps closed. Hunt completion counts no longer appear, and the CISO stopped asking for the heatmap.

This week, pull your last quarter of hunt reports and count how many changed a detection or closed a visibility gap. If the answer is zero, you already know which kind of program you are running, and now you know where to start fixing it.

Frequently asked questions about threat hunting

What does threat hunting actually mean in practice?

The PEAK framework defines threat hunting as any manual or machine-assisted process meant to find security incidents that automated detection missed. Its co-author David Bianco has corrected the emphasis: hunting exists to drive continuous improvement across the security program, and finding incidents is a side effect.

In practice, a hunt is a testable hypothesis about adversary behavior, investigated against your data, that ends in an improved detection or a closed visibility gap. A documented disproof is also a valid endpoint.

How is threat hunting different from alert triage?

Alert triage reacts to alerts a detection has already raised and follows a set workflow to a verdict. A hunt starts earlier, with a hypothesis that an attacker is operating in the environment without triggering an alert, then tests that hypothesis against raw telemetry. The shorthand from the TaHiTI methodology is useful here: anything a tool can do autonomously is not hunting.

How do you measure a hunting program without gaming it?

Use outcome metrics that cover detections created or updated and gaps identified and closed. Track incidents opened after the hunt because of detections the hunt produced. A composite basket is harder to game than a single north-star number, and it keeps Goodhart's Law from converting your program into activity theater. Drop per-hunter hunt counts and standalone ATT&CK coverage percentages, since both reward volume over working detections.

Does a hunt that finds nothing still count?

Yes, provided the hypothesis was testable and the data existed to test it. The TaHiTI methodology counts a disproven hypothesis as a valid outcome, separate from a failed hunt where the data to test the claim never existed. The 2025 joint hunt by CISA and the Coast Guard at a critical infrastructure organization turned up no adversary and still surfaced six security gaps, which is exactly the kind of return a null result should deliver.


About the author

MKMarta K. is a senior detection engineer and incident responder with over eight years of hands-on experience operating and scaling security operations in high-growth SaaS and fintech environments. She started her career as a SOC analyst, working night shifts triaging alerts and investigating suspicious activity across endpoint, identity, and cloud environments. Over time, she moved into detection engineering, where she focused on building and tuning detection pipelines, reducing false positives, and mapping coverage to frameworks like MITRE ATT&CK. Marta has led incident response efforts for ransomware, credential compromise, and insider threat scenarios, and has helped teams transition from reactive alert handling to structured investigation workflows and proactive detection strategies. Her work has included implementing detection-as-code practices, improving alert fidelity, and designing playbooks that actually get used during real incidents. She writes about the reality of running security operations — from alert fatigue and broken escalation paths to what actually works when building detections and responding to incidents under pressure.

Stay sharp on security operations

Practitioner takes on SOC modernization, detection engineering, threat hunting, and more. No fluff. No product pitches.

Most threat hunting programs produce activity theater | Future of SecOps