AI SOC: a working taxonomy for skeptical practitioners

DCDaniel C. · Head of Security Operations
AI in Security Operations·11 min read

Here's the taxonomy I use to separate copilots, triage agents, investigators, detection engineers, and automation platforms, and the validation questions I ask before approving auto-close.

Last quarter, I watched an AI SOC platform close 94% of alerts while the demo dashboard updated in real time. I asked one question: how many of those closures had later been overturned by a genuine finding? The vendor couldn't show me, because the product measured the volume it processed, not the threats it missed.

That's my buying rule for AI SOC tools: don't judge them by alerts processed, auto-close rate, or a standalone accuracy percentage. Judge them by the severity-weighted false negative rate in the alerts they close without human review, measured through blind quality assurance.

Then classify the product before you compare features. Is it a copilot, a Tier-1 triage agent, an investigation tool, a detection-engineering tool, or an automation platform? What can it recommend, close, or act on? And who can inspect, override, and answer for the result when it's wrong?

In brief:

  • "AI SOC" is a broad marketing label covering analyst assistance, Tier-1 triage, investigation support, detection engineering, and workflow automation. Sort products by five jobs and two additional dimensions, not an industry standard: decision authority and accountability.
  • The label reveals neither how the closure decision is reached nor who's authorized to make, override, and answer for it. Safe automation needs deterministic controls around data retrieval, validation, escalation, authorization, auditability, and rollback.
  • Agentic managed detection and response (MDR) and AI SOC software are different purchases. Confirm evidence access, escalation coverage, action authority, audit rights, and contractual commitments before treating either as an accountability transfer.
  • For auto-close, measure severity-weighted confirmed incidents in a blind review of the closed set. Closure rate and throughput measure capacity; they don't establish safety.

The label names a marketing category, not the thing you're buying

Agentic doesn't mean autonomous

For this taxonomy, an AI SOC uses AI agents to triage, investigate, and respond to alerts with limited human involvement. That definition is useful only as a starting point, because it puts a summarization copilot bolted onto a security information and event management (SIEM) platform in the same bucket as a triage agent that only recommends, and the same bucket also includes a platform that quarantines endpoints unattended.

I use agentic to mean the product plans and executes multi-step work, not that it acts autonomously: a platform can be agentic without being autonomous, and can market itself as an AI SOC without being agentic at all. I find AI-assisted or augmented analyst a more practical description of much of the category.

Sort by job, decision authority, and accountability

For this evaluation, I sort products by five jobs and two additional dimensions: decision authority and accountability. First, identify the job: copilot, Tier-1 triage, investigation, detection engineering, or automation.

Second, identify decision authority: advisory only, analyst-approved, policy-gated auto-close, or autonomous action. Third, identify accountability: who can inspect the evidence, who can override the result, who owns escalation, and what obligations apply when the system misses.

Anton Chuvakin has drawn a related distinction in his writing on AI in the SOC: most teams arrive at tools that augment people and sometimes finish work end to end, rather than a wholly AI-run SOC.

Five jobs hide under one label

I start with a four-part market map: AI-extended detection and response (AI-XDR) and copilot platforms, Tier-1 analyst tools, Tier-2/3 analyst tools, and AI automation engineers. I add a fifth for detection engineering, because vendors now offer it as a distinct function. This is an editorial framework, not a verified industry classification, and vendors overlap across it as capabilities shift:

  • Copilots: For this taxonomy, I separate cybersecurity AI assistants from AI SOC agents. They explain alerts and don't own a verdict.
  • Tier-1 triage agents: Dropzone AI, Intezer, Simbian, Qevlar AI, Exaforce, and 7AI currently emphasize automated alert investigation or triage in their market positioning. The job is to investigate supported alert types, recommend or issue a disposition, and attach an auditable evidence chain.
  • Tier-2/3 investigators: Command Zero positions around complex investigation and threat hunting rather than routine Tier-1 alert triage. Different job, different buyer.
  • Detection engineering agents: Prophet Security also addresses Tier-1 triage, but its positioning extends into threat hunting and detection tuning, including a named AI Detection Engineer focused on rule authoring and coverage gaps; I treat it as a cross-category platform rather than a pure Tier-1 tool. The CTI-REALM benchmark evaluates AI agents on end-to-end detection-rule generation from CTI reports, measuring final outcomes alongside trajectory-based rewards for the investigative process.
  • Automation engineers: I place Torq's HyperSOC primarily on the automation-engineering side of this map because its value proposition centers on orchestration and workflows, with AI layered into that operating model.

When a vendor emphasizes multi-agent architecture, ask which supported alert classes the agents handle, what evidence they retrieve, and what decisions or actions they may take without approval.

Who computes the verdict matters more than which model runs

Constrain the LLM, don't just trust it

The production pattern I trust most pairs LLM reasoning with deterministic retrieval, validation, authorization, and escalation controls. The agent may help select investigative steps and organize evidence, but it shouldn't be allowed to close an alert or take action unless required data, policy checks, confidence thresholds, and audit records are present.

Per-rule guidance, asset criticality, identity context, and historical closure rationale can improve investigation quality when retrieved and validated consistently, though auto-close should stay limited to high-confidence false-positive verdicts.

Ask what happens when it's wrong, not just how often

Aggregate performance can hide a consequential miss, and the risk isn't unique to LLMs: a deterministic system fails just as easily when it's tuned wrong, required evidence is stale, or an enrichment API is down. I assess investigation quality using reproducible evidence chains, correct retrieval of relevant context, and appropriate escalation of uncertainty and unavailable data.

The second dimension is whether the system recommends or acts. In the products I've evaluated, designs range from recommendation-only workflows to graduated modes requiring human confirmation and configurable containment that can quarantine endpoints or disable credentials, and some platforms present containment as autonomous or available with one click.

I open every demo with the same question: what actions run before a human sees anything, and can I see what was suppressed and why? If a vendor responds with a single accuracy percentage instead of an answer, it hasn't provided the information needed to evaluate operational safety.

Agentic MDR and AI SOC software are different checks to write

MDR's automation claims still need an audit

The MDR side now uses the same vocabulary. I read the public positioning from Expel, Arctic Wolf, ReliaQuest, and CrowdStrike Falcon Complete as variations on a common model: AI absorbs more of the scale and investigation workload, and it may also handle first-pass decision-making while human analysts remain part of escalation and judgment.

When a provider claims close agreement between automated and human triage decisions, I still want the measurement method and an independent audit; I have not seen one behind the performance figures.

Most MDR offerings still combine automation with human escalation, incident coordination, and service accountability. Verify which decisions are automated, which require provider analyst review, and which remain your responsibility.

Software and service shift the burden differently

Software gives you the evidence chain and the raw queries, along with more of the configuration and governance burden. Confirm whether model-provider costs are bundled, usage-metered, capped, or passed through, and who's responsible for model changes, prompt-injection controls, and rollback.

A managed service can provide an escalation path and a human operating layer, but verify the staffing model, escalation coverage, named points of contact, and response-time commitments before assuming it means one accountable person.

In the MDR agreements I've reviewed, customer obligations, exclusions, and liability limits are typically explicit; provider commitments for missed detections and automated closures may be less specific. Treat those as negotiation points, not assumptions, and ask for audit rights over suppressed alerts before signing.

Three risks every buyer should test

Adversarial input. Prompt injection through attacker-controlled telemetry is a demonstrated risk in LLM-augmented SOC workflows. In one 2026 study, log fields such as URLs, user agents, payloads, DNS queries, and usernames carried adversarial instructions into the model, and summarization was the highest-risk task tested: context-manipulation attacks succeeded 96% of the time without defenses and 38% even under constrained-output defenses. The strongest tested defense cut average injection success from 26.6% to 11.8%, but didn't eliminate it. These are controlled research results, not a measured failure rate across commercial AI SOC products, but the risk is especially relevant to copilots and agents that summarize or reason over raw, attacker-influenced telemetry. A human reviewing the underlying evidence may catch the manipulation where an unattended path turns it into a missed or misrouted alert.

Incomplete context. The most dangerous closure isn't necessarily a model hallucination. It can be a reasonable-looking verdict built on missing asset criticality, stale identity data, incomplete logging, or an unavailable enrichment source. A vendor's claimed 99% accuracy figure isn't usable until you know the denominator, the class balance, and whether the number means precision, recall, or agreement with human analysts. In a queue where only 1% of 10,000 daily alerts are genuine, a system that classifies every alert as benign can still show 99% raw accuracy while missing all 100 real alerts.

Uncontrolled change. Foundation models, integrations, and detection rules can all change during a security contract. Require model-version visibility, change notices, regression testing, rollback procedures, and a clear pricing schedule stating whether inference costs are bundled, metered, capped, or passed through.

What I ask before I write the check

The blind review methodology

My rule for vendor accuracy claims: if a vendor gives me a precision percentage but cannot explain how it was measured, I don't treat the number as usable.

For auto-closed alerts, I want a blind quality-assurance review: an analyst re-investigates a stratified sample without seeing the model's verdict, using only the evidence available at the time of closure. Sample by alert class, severity, telemetry source, and action type, and oversample privileged identities and high-impact categories, since a missed low-severity false positive and a missed privileged-account compromise shouldn't count as equivalent misses.

During shadow mode or a narrowly scoped rollout, review every proposed auto-close; after that, keep a recurring blind sample large enough to catch a meaningful change in the escape rate.

What goes in the RFP

A material confirmed incident in the closed set should trigger a documented review and tighten or suspend auto-close for that category until the failure is diagnosed and the correction validated. Validated, reversible, low-impact actions can run within policy once cleared; escalate anything uncertain, high-impact, or irreversible. Alongside that, four questions go in every request for proposal (RFP):

  • What is your measured false negative rate on auto-closed alerts, how do you measure it, and what were your last three material misses?
  • Which of the five jobs does the agent do, and which tier of my queue does it never touch?
  • What actions run before a human sees anything, where is the decision tree, and can I export every suppressed alert with its reasoning?
  • Which foundation models sit underneath, what changed the last time one of them did, and does my price move when your token bill does?

Start with the pile the AI closed

The adoption-maturity gap is real

AI and ML use is now widespread in security operations, but operational integration is much less mature. SANS reported in 2026 that 79% of SOCs use AI or ML tools, while only 36% have integrated them into a defined workflow; most analysts use AI individually, without organizational structure or systematic validation.

Teams buy anyway to add capacity, and I have also seen security programs disable detection rules because they lack the people to investigate the resulting alerts. Capacity, not risk, is setting coverage in too many shops.

Pick the operating model, then run the pilot

Daylight's operating-model framework argues that decision ownership is the primary evaluation criterion, and distinguishes internal AI SOC software, AI-enabled MDR, and hybrid models by who controls evidence, governance, and escalation.

I weigh that decision on coverage hours, investigation maturity, data quality, regulatory requirements, and appetite for operating the system, not team size alone; a small team with strong data foundations can run the software, and a larger one without investigation maturity may still be better served by the service.

Pick your single noisiest alert category, run one product in shadow mode against your analysts for a month, and re-investigate its auto-closed pile blind. Track the rate of confirmed incidents in that closed set, weighted by potential impact, alongside the analyst time it actually saved. The goal isn't the highest closure percentage; it's a verifiable reduction in low-value work without a matching increase in what gets missed.

Frequently asked questions about AI SOC

What is the difference between an AI SOC platform and agentic MDR?

An AI SOC platform is software your team runs; you own configuration, decision thresholds, and the relationship with the AI vendor. Agentic MDR is a service where the provider runs AI agents on your behalf and its analysts own escalation. Most MDR still combines automation with human accountability, so ask any provider which decisions its AI makes before a human sees the alert.

How do you evaluate an AI SOC tool before buying?

Run it in shadow mode against one bounded, high-noise alert category and compare its proposed dispositions with analyst outcomes. Review every proposed auto-close during initial rollout; after validation, maintain a recurring blind, stratified sample and full review of high-impact cases. Track confirmed incidents and severity-weighted escapes in the closed set, plus the analyst time actually saved. Throughput and closure rate are capacity metrics, not proof of safe automation.

Can AI SOC agents be prompt injected through logs?

Yes. Research has demonstrated that attacker-controlled telemetry can manipulate LLM-based SOC tasks, especially summarization. Instructions embedded in attacker-controlled content can steer an agent toward a benign summary or incorrect evidence. Any agent that reads attacker-controlled content needs deterministic controls around data handling, validation, and action authorization; high-impact or irreversible actions should require explicit human approval, while narrowly scoped, reversible actions can be pre-approved once validated and continuously monitored.

Is the AI SOC market mature enough to buy from in 2026?

Mature enough to run a scoped evaluation, not mature enough to buy as a category. SANS's 2026 survey found 79% of SOCs use AI or ML tools, but only 36% have integrated them into a defined workflow. Start with one defined job, one alert tier, an explicit decision boundary, and measurable closed-set quality.

Will AI SOC tools replace SOC analysts?

The practitioner view I trust keeps humans central to modern security operations. Agents are taking on alert correlation and enrichment, and first-pass verdicts are moving to them as well. People retain business impact and risk prioritization, and incident ownership stays with them. Expect the role to change shape faster than it shrinks.


About the author

DCDaniel C. is a security operations leader with over a decade of experience building and scaling SOC capabilities for cloud-native companies. He has led security teams through multiple stages of growth — from early-stage environments with minimal tooling to mature organizations operating 24/7 security operations with distributed teams. His experience includes designing SOC architectures, evaluating and managing MDR providers, and building internal detection and response capabilities. Daniel has been responsible for vendor selection across SIEM, EDR, and XDR platforms, as well as defining SLAs, response models, and escalation frameworks. He has also worked closely with executive leadership on budgeting, board reporting, and aligning security operations with broader business risk. He writes about the practical decisions security leaders face — including build vs buy tradeoffs, how to evaluate security vendors, and what it actually takes to run an effective security operations function at scale

Stay sharp on security operations

Practitioner takes on SOC modernization, detection engineering, threat hunting, and more. No fluff. No product pitches.

AI SOC: a working taxonomy for skeptical practitioners | Future of SecOps