Select Language

Choose your language

株式会社ヤグラ

Select Language

Choose your language

株式会社ヤグラ

Select Language

Choose your language

Insight

Comparing and Selecting AI SOCs: PoC Evaluation Items and Fillable Checklist

Compare AI SOC with 9 key evaluation criteria and PoC check items. Assess investigation accuracy, permissions, approvals, log loss, failure response, and manual labor hours using our registration-free evaluation CSV. Organize the evidence needed for deployment decisions and internal approvals.

AI SOCとは? 仕組み・従来型SOCとの違い

Updated: September 12, 2026. Added a PoC evaluation table capable of recording data gaps, failures, and manual verification, along with a registration-free CSV.

The number of products and services claiming "AI SOC" is rapidly increasing. Many organizations find that comparing proposals does not clarify what is automated, to what extent, or how these offerings differ from traditional Managed Detection and Response (MDR) and Security Orchestration, Automation, and Response (SOAR) products. The phrase "AI investigates alerts" varies significantly by vendor. Currently, the same terminology is used for everything from simple AI-generated summaries attached to notifications, to systems that construct investigation strategies, gather evidence, and execute response actions for each alert.

In a June 2025 press release, Gartner predicted that over 40% of agentic AI projects will be abandoned by the end of 2027 due to rising costs, unclear business value, or inadequate risk controls. At the same time, "agent-washing" is spreading, with vendors rebranding existing products like AI assistants, RPA, and chatbots as "agentic AI" without providing substantive agent capabilities. Gartner estimates that of the thousands of vendors claiming agentic AI, only about 130 possess genuine capabilities (Gartner, 2025). In security operations, organizations need a framework to verify substance over marketing claims.

This article outlines evaluation criteria for comparing and selecting AI SOCs, detailed metrics to measure during a Proof of Concept (PoC), common failure modes, implementation steps from current-state assessment to KPI review, and key points for internal approvals and executive briefings. The goal is to provide a decision-making framework tailored to your alerts and team structure, rather than ranking specific vendors. For details on how AI SOCs work and how they differ from traditional SOCs, refer to our article "What is an AI SOC? Mechanism, Differences from Traditional SOC, and Benefits Explained."

Why Evaluating "AI SOC" Is Difficult and 9 Evaluation Criteria

Why Evaluation Is Challenging

The industry has not yet standardized the definition of an AI SOC. Generally, it refers to an operational model, product, or service where AI agents replicate the investigation steps of SOC analysts. These agents autonomously triage, analyze, recommend, and execute response actions for alerts from Endpoint Detection and Response (EDR) and Security Information and Event Management (SIEM) platforms. Gartner describes agentic AI as systems that "autonomously plan and act to achieve user-defined goals," predicting that by 2028, at least 15% of day-to-day work decisions will be made autonomously by agentic AI (Gartner, 2024). In Japan, the "AI Guidelines for Business (Version 1.2)" published by the Ministry of Internal Affairs and Communications (MIC) and the Ministry of Economy, Trade and Industry (METI) in March 2026 defines an AI agent as "an AI system that perceives its environment and acts autonomously to achieve specific goals" (MIC & METI, 2026).

The first challenge in evaluation is that the scope of autonomous action varies by provider. Both generating alert summaries and conducting full evidence-based triage followed by containment can be described as "AI-driven investigation." Second, AI performance cannot be evaluated solely in a demo environment. An AI agent's investigation accuracy depends heavily on its data access scope and context comprehension of the customer's environment. Performance in a vendor-curated environment does not guarantee accuracy when processing noisy, real-world alerts. Third, deployment success depends more on human role design and operational processes than the product itself. Gartner data indicates that only 20% of cybersecurity teams report "highly beneficial" outcomes from generative AI use cases, and predicts that through 2027, 90% of successful AI implementations in cybersecurity will be tactical rather than strategic (Gartner, accessed 2026). Rather than pursuing a broad vision, success requires measuring tangible impacts on clearly defined, scoped use cases.

To compare proposals effectively, evaluate "what, to what extent, and how" across the following nine criteria. Differences between adjacent services such as MDR, MSS, and XDR are detailed in "Differences Between MDR, MSS, XDR, and AI SOC: How to Choose Security Operations Services." This article focuses specifically on comparing AI SOC offerings.

Criteria 1–3: Depth of Investigation, Scope of Response, and Product Integration

Depth of investigation is the primary criterion to review. The residual manual workload differs significantly between "notifications with AI summaries" and "structured investigation reports that query multiple data sources to present a timeline and evidence." The former requires analysts to restart investigations from scratch, while the latter shifts the analyst's role to reviewing reports and confirming decisions. When reviewing proposals claiming "AI-driven investigation," request actual investigation reports to verify which data sources are queried. What should be automated during triage is covered in "Alert Triage Automation: AI Agent Investigation Steps and Ensuring Accuracy."

Scope of response defines whether the system supports "recommendations only," "execution after manual approval," or "conditional automated execution," and whether these options are configurable. For high-impact operations like host isolation or account disablement, clarify approval units (per alert vs. per rule), notification methods for approvers, and out-of-hours workflows.

Product integration checks whether the system connects to your active EDR, SIEM, identity providers, email security, and cloud services, and whether integration supports bi-directional actions rather than read-only access. Beyond the number of supported integrations, confirm whether your specific configurations are supported, and the cost and feasibility of custom integrations. For differences compared to SOAR products executing static playbooks, see "Difference Between SOAR and AI SOC: From Playbook Automation to Autonomous Agent Investigations."

Criteria 4–6: Local Context Adaptation, Explainability/Audit Trails, and Human Control

Local context adaptation determines long-term accuracy. Clarify how the system ingests organizational context—such as whether a connection is benign, or whether a server is production or staging—from asset registries, business rules, and historical decisions to apply to subsequent investigations. From a confidentiality perspective, verify that learned data is isolated to your tenant and not commingled with other customers' data.

Explainability and audit trails are critical in regulated industries. The NIST AI Risk Management Framework (AI RMF 1.0), published in January 2023, outlines seven characteristics of trustworthy AI: validity and reliability, safety, security and resilience, accountability and transparency, explainability and interpretability, privacy-enhanced, and fair (with harmful bias managed). It notes that "accountability is presupposed by transparency" (NIST, 2023). For an AI SOC, this means "why an alert was deemed benign" and "which logs were used as evidence" must be recorded in a human-readable format for third-party auditing. Additionally, the NIST Generative AI Profile (NIST AI 600-1) published in July 2024 lists "confabulation" (hallucination) as a key risk (NIST, 2024). A mechanism to cross-reference report evidence against actual logs is essential.

Human control defines the boundaries of AI autonomy and human decision points. The Japanese AI Guidelines for Business (Version 1.2) advise users to "consider involving human judgment at appropriate times rather than relying solely on AI decisions" (MIC & METI, 2026). Verify if approval checkpoints can be configured, if manual overrides are recorded to improve future investigations, and if the scope of automation can be phased or reverted.

Criteria 7–9: Security and Data Handling, Operational Support, and Cost Structure

Security and data handling requires confirming where logs and investigation data are stored, retention periods, encryption, access controls, and whether your data is used for model training. Additionally, AI agents themselves can be attack targets. Gartner predicted in an August 2026 press release that by 2029, over half of successful cyberattacks on AI agents will exploit weak access controls and prompt injection (malicious instructions embedded in AI inputs) (Gartner, 2026). Because AI SOCs hold response privileges, verify that the vendor minimizes privileges assigned to agents and designs the system to resist prompt injections from external data sources like alert payloads or log files.

Operational support defines how the vendor handles system failures or ambiguous AI decisions, the availability of local-language support, support hours, and who manages detection rule tuning and integration maintenance. Since an AI SOC requires continuous tuning as your environment changes, vendor alignment with your operational hours is a key differentiator.

Cost structure details the billing metrics (endpoints, log volume, alert volume, users), cost predictability during alert surges, initial integration costs, and terms for data return or deletion upon contract termination. Compare the total cost of ownership, including the manual effort remaining on your team (reviewing reports, handling approvals, managing exceptions) alongside product fees.

Specific Metrics to Verify in a PoC

While proposals clarify architecture, AI investigation accuracy and context adaptation must be verified through hands-on testing. When designing a PoC, document the duration (typically several weeks to two months), alert scope, success criteria, and evaluation owners before starting. Setting criteria beforehand prevents shifting targets post-testing. Measure the following seven metrics during a PoC:

1. Historical alert replication. Select dozens to hundreds of historical alerts from your environment, including true positives (actual attacks or policy violations), false positives, and ambiguous cases. Run them through the AI to compare its conclusions with historical human decisions. This measures investigation quality on your actual data rather than vendor demos. Including real historical incidents is also critical for measuring false negatives.

2. False negative and false positive rates. Cross-reference AI determinations (Benign / Action Required / Review Needed) against human decisions. Separately calculate true positives marked "Benign" (false negatives) and benign alerts marked "Action Required" (false positives). In SOC operations, false negatives can lead to security breaches, while false positives increase analyst workload; they should not be weighted equally. Note the distribution of confidence levels and "Review Needed" classifications to verify if the AI correctly identifies its own uncertainty.

3. Average investigation time and throughput. Measure time elapsed from alert ingestion to report generation, time required for analysts to verify and confirm decisions, and processing capacity during alert volume spikes. To maintain consistency, we recommend using standardized definitions for MTTD (Mean Time to Detect) and MTTR (Mean Time to Respond) as outlined in "What are MTTR and MTTD? Practical SOC KPI Design and Improvement."

4. Report quality. Ensure multiple analysts evaluate whether reports document what was queried, the evidence used, and the reasoning in a chronological format. Verify if affected hosts, accounts, and destinations are structured, if attack techniques are mapped to common frameworks like MITRE ATT&CK, and if recommended actions and uncertainties are clear. Randomly cross-reference report evidence against raw logs to check for confabulation.

5. Escalation logic. Verify the criteria for human intervention, whether escalation severity aligns with your internal standards, and whether notification delivery times during off-hours meet operational SLA requirements. It is critical to monitor both "under-automation" (excessive human escalation) and "over-automation" (failure to escalate when needed).

6. Localization and language support. Confirm that reports, consoles, and support portals are available in your team's operating language. Ensure the system correctly parses localized hostnames, usernames, internal ticketing systems, and policy documents, and accounts for local business hours and holidays. Reports must be clear enough to present to management or auditors without translation delays.

7. Audit logs. Verify that all AI queries, decisions, and response actions—as well as human approvals or rejections—are recorded in a tamper-resistant format, retained for the required duration, and exportable for auditors. The AI Guidelines for Business (Version 1.2) recommend recording and storing "logs of reasoning processes and decision bases" within reasonable limits (MIC & METI, 2026). For organizations in regulated sectors like finance or defense supply chains, these logs form the foundation of compliance reporting.

PoC Evaluation Table: Tracking Ambiguity and Residual Workloads

Once evaluation criteria are set, document the input parameters, AI outcomes, human conclusions, and residual tasks for each alert. Below is an operational evaluation table designed by Yagura. This is a practical framework, not a specific product's performance data or compliance audit standard. Establish your success thresholds before starting the PoC, aligned with your operations and risk tolerance.

Testing Scenarios Beyond Nominal Cases

Evaluation Scenario

Expected Behavior

Required Records

Distinguishing attacks from benign operations

For similar alert types, verifies raw logs and business context before making a decision.

Initial AI determination, final human determination, evidence logs

Incomplete or missing logs

Does not treat lack of evidence as proof of safety. Clearly states what data is missing.

Scope of retrieved data, period of missing data, reason for suspension/escalation

Insufficient investigation privileges

Does not treat retrieval failures as successes. Flags missing permissions and requests human review.

API/connection failures, affected decisions, handoff to human analyst

Conflicting evidence

Does not select only favorable evidence. Highlights contradictions and lists manual verification steps.

Conflicting log sources, unresolved queries, human verification results

Response actions requiring manual approval

Handles containment or isolation actions in strict accordance with configured approval policies.

Proposed action, approval/rejection status, execution details and results

Interrupted or failed investigations/responses

Distinguishes completed steps from incomplete ones; never reports a failed action as completed.

Executed, failed, and unexecuted actions; recovery/handoff results

Manual override of AI conclusions

Maintains clear traceability of the original AI reasoning and the subsequent human override.

Reason for override, analyst name, additional data reviewed, and time spent

Run scenarios involving high-risk response actions in non-production environments with restricted targets and permissions. If a product lacks execution capabilities, log it as "Not Supported" and evaluate the recommended actions and the remaining manual steps. Ensure unsupported capabilities are tracked separately from operational failures.

Download PoC Evaluation Table (CSV, No Registration Required)

How to Use the Evaluation Table

  1. Standardize target products, evaluation periods, alert selection criteria, available log sources, permissions, and approval workflows. Track alerts selected from natural distributions separately from test cases designed to evaluate data gaps or system failures.

  2. Record each evaluation run as a single row. If re-running the same alert, assign a new test ID to avoid mixing initial and subsequent results. Use the "Input Template" row for data entry, and refer to the "Example Case" row to see how to document results. Do not include example rows in your final performance metrics.

  3. Retain the initial AI determination even if it is later overridden by a human. For alerts where a final human decision cannot yet be made, log them as "Undetermined" and track their count separately rather than omitting them from accuracy denominators.

  4. Include time zones in all timestamps. Separately record timestamps for report generation, human validation, approval of response actions, and actual execution completion. For cases where response actions were unnecessary or unsupported, document the reasons and leave completion fields blank.

  5. Measure both AI processing times and human review/investigation times. Keep queuing delays separate from active working times.

Standardizing Metrics for Aggregation

For the investigation completion rate, define the total alert pool size upfront. Explicitly report the count of unstarted, in-progress, suspended, or failed cases. Alerts suspended due to insufficient evidence should be tracked as a correct behavior (suspension) rather than aggregated with nominal completions.

For false negatives, report the count of alerts confirmed as "Action Required" by a human but classified as "Benign" by the AI, alongside the total true positive volume. For false positives, report the count of alerts confirmed as "Benign" by a human but classified as "Action Required" by the AI, alongside the total benign volume. Do not force "Review Needed" or undetermined cases into binary categories. For small sample sizes, report absolute counts alongside percentages.

For time metrics, break down intervals such as "Alert Ingestion to Report Completion," "Report Completion to Human Validation," and "Response Approval to Execution." Report completed and pending counts, and compare medians and averages under identical evaluation conditions. Document total human effort for validation and supplemental investigation to ensure fast report generation does not shift the workload to manual verification.

Finally, map outcomes against your pre-defined success criteria using classifications such as "Pass, Fail, Untested, or Not Supported," and assign owners and deadlines for any outstanding issues. If you are evaluating Yagura AI SOC, you can use this table to discuss your specific validation requirements with us. Documenting your target products, problematic alerts, accessible logs, and intended analyst roles will help focus the evaluation scope.

Common Failure Modes

Relying solely on vendor demos. Demos run on vendor-curated datasets show systems under optimal conditions. Demos evaluate UI usability and report formats, not accuracy against your real-world alerts. Contracting without historical alert replication often leads to high error rates in production, forcing analysts to manually re-verify every alert.

Using overly simplified test cases. Restricting a PoC to clean, unambiguous alerts to reduce testing overhead fails to measure how the system handles noise, duplication, and missing logs—which constitute the majority of production volume. Conversely, routing all production alerts immediately can overwhelm evaluators, making detailed validation impossible. Test with a statistically representative sample of varying alert types and qualities.

Unclear human role definition. Deploying an AI SOC without defining who reviews alerts, who authorizes containment, and who holds override authority leads to either "automation bias" (blindly trusting AI decisions) or "systemic distrust" (manually re-investigating every alert). The NIST Generative AI Profile warns of automation bias and over-reliance as significant human-AI configuration risks (NIST, 2024). Mitigation requires designing "human-in-the-loop" validation processes during the PoC phase, mimicking production roles.

Lacking clear success criteria and KPIs. Initiating deployments based on vague performance impressions makes it difficult to justify investment during procurement or post-deployment reviews. A September 2026 Gartner survey notes that only 22% of organizations have successfully scaled AI across business units, highlighting that organizations tracking expenditures against specific business outcomes are better positioned to protect investments and redirect resources away from non-performing initiatives (Gartner, 2026). PoC success criteria should translate directly into operational KPIs.

Structured Deployment Framework

Aligning evaluation criteria and PoC metrics into an implementation roadmap yields a five-phase process:

  1. Current-State Assessment: Document and quantify baseline alert volumes, triage rates, MTTD/MTTR, manual effort, active EDR/SIEM configurations, and regulatory/audit requirements.

  2. Proof of Concept: Test the seven key metrics against historical alerts and a scoped selection of production alerts, measuring outcomes against predefined success thresholds.

  3. Shadow Operations: Run the AI SOC in parallel with live production alerts. The AI generates investigations while human analysts make final decisions. Measure decision alignment and time-to-conclusion differences.

  4. Phased Automation: Automate low-risk response actions (e.g., log collection, ticket creation, stakeholder notifications) first. Phase in high-impact actions like host isolation with mandatory manual approvals.

  5. KPI Review: Conduct monthly and quarterly reviews of triage rates, false positive/negative rates, operational response times, and manual workloads to refine automation boundaries and approval rules.

In the current-state assessment, identify "what percentage of generated alerts are currently investigated." In many organizations, resource constraints mean only high-severity alerts receive manual triage, while others are auto-closed or expire. This baseline is essential to measure how the triage rate changes post-deployment.

Shadow operations act as a safety gate between the PoC and production. Comparing AI determinations against live analyst decisions helps identify which alert categories are stable enough to transition to automated workflows.

During phased automation, establish "AI investigates, human approves" as the baseline. As decision alignment stabilizes for specific alert categories, transition those approvals to post-execution reviews. Documenting the criteria for expanding or restricting automation scope simplifies reporting to compliance and executive stakeholders.

Internal Approvals and Executive Briefings

When presenting an AI SOC business case to executives, structure materials to address the three primary reasons agentic AI projects fail: cost escalation, unclear business value, and inadequate risk controls (Gartner, 2025). The April 2026 Gartner Hype Cycle for Agentic AI notes that while only 17% of organizations have deployed AI agents, over 60% plan to do so within two years, driving increased corporate focus on accountability, control, and economic viability (Gartner, 2026). Addressing these areas directly helps secure executive alignment.

First, define business value as "improvements in alert triage coverage and time-to-containment" rather than head-count reduction. Use baseline metrics and PoC results to show how the system reduces exposure window risks. Frame manual time savings as an opportunity to redirect analyst focus toward proactive threat hunting and detection engineering.

Second, present costs comprehensively. Include product fees, implementation effort, and the "cost of inaction" (unresolved off-hours alerts, analyst burnout, recruitment, and retention costs). Detail the pricing model, billing metrics, and cost caps during alert spikes to address budget predictability concerns.

Third, address risk controls using industry frameworks. The NIST AI RMF 1.0 outlines core functions—Govern, Map, Measure, and Manage—as a voluntary framework for AI risk management (NIST, 2023). Locally, the AI Guidelines for Business (Version 1.2) outline ten guiding principles, including human-centricity, safety, transparency, and accountability, advising users to ensure safe deployment, robust security, and stakeholder transparency (MIC & METI, 2026). Furthermore, the ISO/IEC 42001:2023 standard specifies requirements for establishing, implementing, and continually improving an Artificial Intelligence Management System (AIMS), emphasizing traceability and accountability (ISO, 2023). Framing the AI SOC deployment as an extension of these compliance frameworks—supported by configured approval gates, structured audit logs, and scheduled reviews—helps address governance concerns from risk and compliance teams.

Yagura AI SOC Capabilities

The evaluation and PoC steps outlined above apply to any AI SOC selection process. For comparison, the following details how Yagura AI SOC aligns with these criteria. Yagura AI SOC is a proprietary service featuring AI agents designed to replicate the workflows of tier-3 analysts to investigate and respond to EDR and SIEM alerts 24/7. It integrates with over 100 security products, including Microsoft Sentinel and leading EDR/identity solutions (Criteria 1–3). The system utilizes context memory to adapt to your specific organizational environment, improving investigation precision over time (Criterion 4).

Yagura AI SOC focuses on reducing operational workloads through automated, context-aware alert analysis. Its impact can be measured using the metrics defined here, such as average investigation time, alert triage coverage, and analyst time savings. We encourage prospective customers to run a PoC and shadow operations using actual historical and production alerts to measure performance in their own environments. Verifying explainability, audit trail robustness, human-in-the-loop controls, and data privacy policies against standard benchmarks ensures long-term operational trust.

Summary

The capabilities behind the term "AI SOC" vary significantly by vendor, requiring validation beyond marketing proposals and demos. When selecting a solution, structure your evaluation across nine areas: depth of investigation, scope of response, integration depth, context adaptation, explainability/audit trails, human control, security, support operations, and cost predictability. Measure performance in a PoC using actual historical alerts to track false negatives/positives, processing times, report quality, escalation logic, localization, and audit trail completeness. Implementing a phased deployment—from baseline assessment through PoC, shadow operations, gradual automation, and structured KPI reviews—while defining analyst roles early is the most reliable path to operational success. Finally, address executive concerns regarding value, cost, and governance by using empirical PoC data mapped to established AI risk management frameworks.

Related Service: For details on our autonomous AI agent service that investigates and responds to EDR and SIEM alerts 24/7, visit the Yagura AI SOC page.

References and Sources

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化や

サービスの概要資料についてお送りいたします。

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化やサービスの概要資料についてお送りいたします。

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化やサービスの概要資料についてお送りいたします。