VYPR
researchPublished Aug 26, 2026· 1 source

Cisco Talos Evaluates AI Models for SOC Efficiency, Finds Cost and Consistency Key

Cisco Talos research reveals that selecting AI models for Security Operations Centers requires balancing effectiveness, cost, and consistency, as higher reasoning effort doesn't always yield better results.

Choosing the right artificial intelligence (AI) model for security operations center (SOC) and digital forensics and incident response (DFIR) tasks is a complex decision, far removed from simply picking the highest-scoring option. Cisco Talos conducted an extensive evaluation of 66 different model and reasoning combinations from leading providers Anthropic and OpenAI, specifically testing their efficacy in log analysis. The study aimed to determine if a clear winner could be identified for these critical security workflows.

Instead of a single standout model, the research highlighted a repeatable methodology for organizations to conduct their own evaluations. A key finding was that increasing the "reasoning effort" – a setting that influences how deeply an AI model analyzes a problem – did not consistently improve results. In many instances, higher reasoning effort led to increased costs without a corresponding improvement in accuracy, and sometimes even resulted in lower scores.

Consistency emerged as a paramount factor for SOC and DFIR operations. Even models that demonstrated strong median performance across multiple tests could produce significantly weaker individual results. In a security context, unpredictable performance can lead to critical errors, such as missed threats (false negatives) or misidentified incidents (false positives), making consistent reliability a crucial metric.

The experiment involved a tool-assisted log-review task where human reviewers, using AI models, had to determine if a given dataset was real or synthetically generated. The dataset mimicked real-world enterprise logs, including network, perimeter, and endpoint telemetry. Each reviewer, acting as one of four distinct personas (Threat Hunter, Detection Engineer, Network Forensics Analyst, and Host/EDR Analyst), analyzed the logs and assigned a confidence score.

Beyond the accuracy score, Talos meticulously measured cost and time for each condition. Cost was calculated based on public API rates, accounting for all attempts, including retries. Time measured the total elapsed duration for completing a set of analyses. The "downside score consistency" metric was also introduced, defined as the difference between a panel's median score and its lowest score, with smaller values indicating greater reliability.

The findings underscore that the optimal AI model is not a one-size-fits-all solution. Organizations must consider a multifaceted approach, weighing the investigative quality against the acceptable tolerance for cost, processing speed, and the potential for inconsistent outcomes. The "best" model is the one that best fits the specific needs and constraints of a particular security team's workflow.

This research provides a valuable framework for cybersecurity professionals navigating the rapidly evolving landscape of AI integration. By focusing on a balanced approach that prioritizes cost-effectiveness and consistent performance alongside analytical accuracy, security teams can make more informed decisions about deploying AI tools to enhance their defensive capabilities.

Synthesized by Vypr AI