Cost-Aware Security AI: The Metric That Kills Demo Agents
Security AI benchmarks have been measuring peak ability, not operational cost. A new cost-aware framework reveals that many top-performing agents are too expensive for real-world use, forcing a rethink of procurement and deployment.
- New arXiv preprint (2607.15263v1) proposes cost-aware evaluation for security agents, measuring success per dollar, not just success rate.
- Offensive Cybench challenges and defensive Splunk BOTS v1 are used to show that top agents often cost 10x more per task than cheaper alternatives with similar success.
- The paper's framework gives buyers a new metric to compare agents, exposing which vendors deliver real value and which rely on brute-force inference.
Why Does Cost-Aware Evaluation Matter More Than Peak Success Rate?
According to the arXiv preprint 2607.15263v1, security-agent evaluations "commonly measure peak offensive capability under generous inference budgets." The authors argue this is "useful but incomplete" because in operational security, "every reasoning step, tool call, telemetry query, and enrichment request consumes budget." This is a direct challenge to the current practice of running agents with unlimited inference budgets on CTF challenges. The paper's framework introduces cost-success curves that plot success rate against cumulative cost, revealing that many agents that look dominant at high budgets are outperformed by cheaper alternatives under realistic constraints. For example, an agent that solves 90% of Cybench challenges at a cost of $500 per task is less valuable than one that solves 80% at $20 per task, yet current benchmarks would rank the first agent as superior.
Who Benefits From This New Framework?
Buyers of security AI – SOC managers, CISOs, and procurement teams – gain the most. The framework provides a transparent, apples-to-apples comparison of total cost of ownership (TCO) for agents. According to the paper's evaluation on Splunk BOTS v1, a defensive agent that queries telemetry aggressively may achieve high detection rates but at a cost that makes it unsuitable for 24/7 monitoring. The framework exposes this tradeoff. Vendors who optimize for cost-efficiency – likely smaller players with leaner models or smarter orchestration – can now demonstrate value against incumbents who rely on expensive foundation models. The losers are vendors who have built their pitch on raw success rate without disclosing inference cost. They will now face tough procurement questions.

How Does This Change Procurement for Security Teams?
Procurement processes today often ask for a demo success rate on a benchmark. The new framework changes the question from "How many challenges can your agent solve?" to "What is the cost per solved challenge under realistic budgets?" According to Splunk's BOTS v1 documentation, the dataset simulates realistic SOC telemetry, making it a better proxy for operational cost than synthetic CTF challenges. The paper's methodology allows a buyer to say: "I have a budget of $10,000 per month for agent inference. Which agent maximizes my success rate within that budget?" This is a fundamentally different, more actionable question. Teams that adopt this framework will avoid overpaying for agents that look good in demos but bleed budget in production.
What Are the Operational Tradeoffs Between Offensive and Defensive Agents?
The paper evaluates both offensive (Cybench) and defensive (Splunk BOTS v1) agents, revealing asymmetric cost structures. Offensive agents often require expensive, multi-step reasoning for exploit development, making them cost-sensitive per attempt. Defensive agents, by contrast, may run continuously, so even small per-query costs accumulate rapidly. The tradeoff: a defensive agent that queries 10x more telemetry sources may catch more threats but cost 50x more. The framework forces teams to decide where to deploy budget: high-cost/high-reward offensive agents for critical targets, or broader, cheaper coverage for day-to-day monitoring. The paper's cost-success curves make this tradeoff explicit.
Which Vendors Should Be Most Concerned?
| Vendor | Typical Approach | Cost Risk | Verdict |
|---|---|---|---|
| Large model providers (e.g., OpenAI, Anthropic) | Use frontier models for agent reasoning | High per-token cost; high success rate at high budget | Need to offer cost-efficient tiers or risk losing procurement battles |
| Specialized security AI startups (e.g., Dropzone AI, Resistant AI) | Use smaller, fine-tuned models with optimized toolchains | Lower per-task cost; may sacrifice peak success rate | Positioned to win on cost-efficiency |
| Incumbent SIEM vendors (e.g., Splunk, Microsoft, SentinelOne) | Embed agents into existing platforms | Cost hidden in platform licensing; agents may inflate cloud compute bills | Must provide cost transparency or face customer pushback |
| Verdict | Specialized security AI startups with cost-optimized agents gain a competitive edge. Large model providers must adapt or lose cost-sensitive buyers. | ||
My thesis: The security AI industry has been running a beauty pageant, not a utility test. The arXiv preprint 2607.15263v1 is the first serious attempt to measure what matters in production: cost-adjusted success. In the short term, this framework will primarily be used by sophisticated buyers – large SOCs with dedicated evaluation teams. In the long term, I expect it to become a standard procurement criterion, similar to how latency and throughput are standard for API pricing. The winners will be vendors who can demonstrate high success at low cost – likely those using smaller, fine-tuned models and efficient orchestration. The losers will be those who have relied on big, expensive models and generous inference budgets to pad their benchmark scores. I predict that within 12 months, at least one major security vendor (likely CrowdStrike or SentinelOne) will publish cost-success curves for its own agents, either voluntarily or under procurement pressure.
- By Q3 2027, at least one major SOC platform (e.g., Splunk or Microsoft) will incorporate cost-success curves into its product evaluation documentation.
- A specialized security AI startup will publish a cost-success benchmark showing its agent outperforming OpenAI's and Anthropic's agents on the Cybench cost-adjusted metric within 6 months.
- The next version of the Cybench benchmark will include a mandatory cost-reporting field for all submissions, driven by this paper's influence.
- 2019-2023Early security agent benchmarks
Focus on CTF success rate; no cost measurement.
- 2024-2025Splunk BOTS v1 and Cybench become standard
Benchmarks established but still ignore cost.
- July 2026arXiv 2607.15263v1 published
Proposes cost-aware evaluation framework for security agents.
- 2026-2027 (projected)Procurement teams adopt cost-success curves
Vendors adapt to new metric; specialized startups gain edge.
Estimated Cost-Success Curve (Cybench Offensive Agent)
Estimated Cost-Success Curve (Cybench Offensive Agent)
X-axis: Cumulative Inference Cost (USD), Y-axis: Success Rate (%)
Agent A (Large Model): Reaches 90% at $500, then plateaus.
Agent B (Specialized): Reaches 80% at $40, then plateaus.
Note: Data is illustrative based on paper's methodology.
- The cost-aware framework from arXiv 2607.15263v1 is the first to measure security agents on cost-adjusted success, not just raw ability.
- Buyers gain a procurement lever to compare TCO; vendors face a new metric that rewards efficiency, not brute force.
- Specialized security AI startups are positioned to win; large model providers must adapt or lose cost-sensitive accounts.
- The framework exposes that many top-performing agents are too expensive for real-world SOC operations.
- Adoption of cost-success curves will become a standard procurement criterion within 12 months.
Source and attribution
arXiv
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Discussion
Add a comment