SecOps leaders ask me some version of the same question midway through an AI SOC evaluation: is the model any good? I've spent this year reading vendor engineering blogs, trust pages, and subprocessor lists, and the question can't be answered as asked. The model usually reduces to one of three types, each with a different job, though the landing page collapses them into one adjective.
A supervised classifier learns from labeled alerts and scores on precision and recall. An anomaly or UEBA (user and entity behavior analytics) model learns normal from mostly unlabeled data, and scores better against historical incidents than lab metrics alone.
A large language model (LLM) plans an investigation and writes the narrative, and if it renders a categorical verdict, that verdict still scores on precision and recall, but only once you also test for run-to-run variance. Which type is doing the work determines what the tool can miss, and why practitioners rank machine learning-powered detection methods near the bottom of what works.
In Brief:
- Supervised classifiers, anomaly baselines, and LLM reasoning layers need different data and fail in different directions under one AI-powered detection label.
- A supervised classifier is only as good as its labels, and labels move: inconsistent definitions of benign activity can dominate its errors.
- UEBA models often lack a fixed ground truth for normal, so a baseline window that includes a breach or reorg can distort what it treats as expected.
- LLM layers plan and narrate. Precision and recall still apply to a categorical verdict, but you also need stability checks, since the model can change through vendor updates unless your contract requires notice.
AI-powered is doing three different jobs, and vendors rarely say which
Broad definitions of machine learning are true and not useful for a buyer: they put a labeled malware classifier, an unsupervised behavioral baseline, and a foundation model trained by next-token prediction under one word. LLMs count too, since pretraining commonly uses self-supervised objectives before supervised fine-tuning. But being machine learning doesn't make something a detector: general pretraining isn't the same as training a classifier on a curated, human-labeled attack dataset.
Products doing materially different jobs still get marketed under the same label, so treat any vendor's public disclosure as a starting point rather than the final word. Per public material, current as of this piece, Exaforce distinguishes among entity resolution, behavioral anomaly scoring, and LLM-based reasoning, and Expel documents a supervised classifier for high-confidence benign phishing alongside an LLM that drafts closure comments after an analyst has already called an alert benign.
Disclosure is uneven elsewhere: Dropzone AI names third-party foundation-model providers, Arctic Wolf has announced an Anthropic collaboration for autonomous-SOC research, ReliaQuest has named both OpenAI and Anthropic while brokering across models, Prophet Security and 7AI didn't name a provider or classifier, and CrowdStrike describes an LLM classifier trained on Falcon Complete triage decisions. Ask each vendor directly for its current model inventory before you trust the label.
Supervised classifiers: what they need, and what they can't tell you
A supervised classifier needs labeled examples, a precision/recall target, and a test set that looks like next month, not last month. Labels usually come from analysts dispositioning alerts as malicious or benign at substantial weekly volume, and they drift under the model. A classifier can perform well overall while concentrating its mistakes in a category analysts define inconsistently, and malware labels show similar instability as samples get rescanned.
The classifier is weakest on a class it never saw labeled, and a lab score can't tell you its real-world hit rate: scores from biased evaluations can fall sharply once spatial and temporal bias are removed, and base rates make the gap worse. A detector that scores ten million benign events a day and falsely flags just 1% of them still produces 100,000 false positives a day before any suppression, and that volume can become the limiting factor for an intrusion detector.
Vendor headline figures need the same translation: an accuracy claim may represent agreement with the vendor's own analysts, not independent ground truth, so ask which set the number came from and how it was split. That skepticism is backed by practitioners themselves: 64% cite high false-positive rates and 61% cite accuracy issues as the top challenges with vendor-provided detection tools, per SANS's 2025 Detection Engineering Survey.
Anomaly and UEBA models: no labels, no ground truth, a different failure mode
An anomaly model learns structure from mostly unlabeled data, though production systems often layer in rules or supervised feedback. A common UEBA design builds an entity's baseline from its own history, its peers, and the organization as a whole, then emits a relative risk score.
Baseline windows vary by signal: a familiar internet service provider (ISP) may be learned over a shorter period than a user's country-level access pattern. Some platforms renormalize older scores when more extreme activity appears, so the same displayed score at two different times may not mean the same deviation. A mature system may weigh peer context, asset criticality, or known-bad indicators too, but there's rarely a single ground truth to check it against.
If the baseline period includes a breach, the model learns the attacker's behavior as expected; mass onboarding or an Active Directory (AD) restructure during that window can distort the baseline the same way. Impossible travel is a familiar false-positive source, and tuning around known virtual private network (VPN) egress can reduce that noise, though blanket suppression of failed logins can hide real credential abuse.
Opaque scores create another problem: without the observations behind a high risk score, analysts can't verify the result or improve the rule, so confidence falls and tuning gets less consistent. Treat anomaly-rule tuning as recurring maintenance, not a one-time task.
LLM-based reasoning: narration, not a verdict you can benchmark the same way
The LLM layer does the parts of an investigation that look like reading and writing. In a typical agentic workflow, the model interprets an incoming alert, plans investigation steps, selects tools, analyzes their responses, and repeats the loop until it has enough data. A triage agent may then evaluate the assembled evidence and propose a verdict.
Either way, the LLM is reasoning over an alert that already exists rather than generating the detection itself, though some platforms also use statistical or rules-based analytics upstream to create it. That reasoning and narrative-writing, the parts a human would otherwise write, is what most AI SOC vendors sell under the AI-powered label.
Score an LLM's verdict the way you'd score any classifier, precision and recall against a labeled reference set, but only once you account for two failure modes unique to language models: run-to-run variance, since some deployed systems produce different answers to the same alert even at temperature 0, and silent drift, since provider-side updates can change output quality without changing the exposed model name. The most dangerous version of that unreliability is when the model papers over missing evidence with a plausible story, treating a failed query as an all-clear and closing a real incident with unwarranted confidence.
For a proof of concept, score verdicts against a senior analyst panel with precision, recall, and Cohen's kappa together (kappa alone can be distorted by rare true positives), run enough repeat trials to see whether the variance is small or wide, and treat any benchmark result you're shown as model- and scenario-specific rather than stable evidence of accuracy.
Why practitioners rate machine learning-powered as one of the least effective methods
Machine learning-powered detection ranks near the bottom of what practitioners find effective, at just 14.2%, per SANS's 2025 Detection Engineering Survey of 264 practitioners. Behavior-based detection led that same survey at 67%, threat intelligence-driven and correlation approaches followed at 43% each, and signature-based detection still outscored machine learning at 41.1%. Yet 45% of organizations already use AI for detection, and 88% expect it to matter within three years.
Model architecture can be treated as vendor plumbing, since nobody buys a diagram and buyers care whether the queue shrinks and attacks get caught, but architecture still decides which outcomes you can measure, and that measurement problem is widespread: 49% of organizations report difficulty evaluating detection engineering effectiveness at all, per the same SANS survey. A signature is inspectable and a behavior-based rule encodes analyst knowledge you can read; an anomaly score is often opaque unless the vendor exposes contributing features, a classifier's labels may come from the vendor's own investigations, and an LLM's answer can vary by run. Baselines built during a merger are the version I see most often: six months later, the rules above threshold are muted and the team is back to hand-written Kusto Query Language (KQL).
Ask which product generated the alert, not just which one scored it
Take the common claim that a platform deploys specialized AI agents for alert triage, threat investigation, and incident response. That product may also list Detection as a capability, described as triaging and classifying alerts from existing tools before they reach the team. Read both claims together and the model type resolves: an LLM agent reasoning over alerts that your own endpoint detection and response (EDR) and security information and event management (SIEM) tools already produced.
When a vendor discloses no labeled classifier, baseline model, or foundation-model provider, that's a signal, not proof, that an LLM agent is doing the reasoning, so ask for the model inventory directly. This pattern recurs across the vendors I reviewed, though some are explicit their AI works only on first-party detections.
For each verdict, ask which product generated the underlying alert and which model produced the verdict. A classifier requires precision and recall on a temporally split test set, plus an account of how the labels were obtained. For an anomaly model, ask about the baseline window and how it handles service accounts, since blanket exclusion can hide the exact accounts attackers target most.
An LLM evaluation should identify the foundation model and the contract's model-swap notification terms, then measure run-to-run variance on your own alerts. That contract language matters more than it sounds: among organizations that reported an AI-related breach, 97% said they lacked proper AI access controls, according to IBM's 2025 Cost of a Data Breach Report. Put the allocation of models to verdicts on your own architecture diagram before any vendor puts it on theirs.
Frequently asked questions about ML in AI SOC tools
AI-powered mechanisms have distinct data requirements and evaluation methods.
What machine learning models do AI SOC tools actually use?
AI SOC tools use supervised classifiers trained on labeled alerts, unsupervised anomaly and UEBA models scored against a learned baseline, or both, and some add LLM reasoning layers that plan enrichment and write investigation narratives. Public material shows several combinations: Expel pairs a classifier with an LLM; Exaforce combines a statistical anomaly model with graph-based entity resolution and LLM reasoning; and Dropzone AI, Prophet Security, and 7AI disclose LLM agents without a separately disclosed classifier. Arctic Wolf has announced an Anthropic collaboration for autonomous-SOC research, ReliaQuest has named both OpenAI and Anthropic while brokering across models, and CrowdStrike describes an LLM triage classifier for first-party detections. Most of these tools triage alerts your existing EDR or SIEM already generated, so confirm each product's actual data inputs.
Is an LLM the same thing as machine learning in this context?
An LLM is a kind of machine learning. Foundation models are commonly pretrained with self-supervised objectives, then fine-tuned with supervised or preference data, but that general pretraining still isn't the same as training a security classifier on a curated, human-labeled attack dataset. In a SOC tool, the LLM interprets alerts, calls tools, synthesizes evidence, and proposes a verdict with a written rationale. Evaluate it on agreement, consistency, and run-to-run stability, and add precision and recall whenever the output is a categorical verdict.
How do you evaluate a supervised classifier versus an anomaly-detection model?
Score a classifier with precision and recall on a test set split by time, not random cross-validation, since temporal bias can overstate post-deployment performance. Run small, single-technique adversary emulation tests mapped to the MITRE ATT&CK framework, define what counts as a successful detection first, and measure per-technique recall in your own environment. For an anomaly model, wait until the baseline window has fully elapsed, then measure false positive volume and analyst actionability, and verify how it treats service accounts and automation identities, since excluding them by default can create a blind spot rather than close one. Ask the vendor, for both model types, how it gets retrained as your environment changes.
Why do AI SOC vendors avoid naming their underlying model architecture?
Many depend on third-party foundation models, so naming one specifically can commit the vendor to something it doesn't fully control. Public disclosures vary: some vendors name multiple commercial model providers, others name none. Foundation-model providers control their own deprecation schedules, so buyers should ask whether their contract requires advance notice and regression testing before any underlying model changes. Whatever the reason for staying vague, the buyer is left responsible for catching a silent model swap unless the contract says otherwise.