A managed SOC should reduce the time between an attacker doing something important and your team containing it. Tool licences, dashboards and alert counts are inputs, not outcomes. The right provider can show which telemetry it needs, which attacker behaviours it detects, how an analyst investigates an alert, who is allowed to respond, and how performance will be measured with evidence rather than averages.
Decide which operating model you need
| Model | Provider owns | Your team owns | Best fit |
|---|---|---|---|
| SOC build | Architecture, onboarding, rules and handover | Daily operations after transition | Teams creating an internal SOC |
| Co-managed SOC | Selected shifts, engineering, investigations or escalation | Shared triage and response | Existing teams with coverage or skill gaps |
| Managed SOC | Monitoring, triage, investigation, reporting and agreed response | Business decisions and retained accountabilities | Teams outsourcing most daily operations |
| MDR | Endpoint-led detection, investigation and response | Identity, cloud, app and broader SIEM coverage unless included | Fast endpoint response with a narrower initial scope |
Ask every bidder to label its offer clearly. 'Managed SOC', 'MSSP', 'MDR' and 'SIEM monitoring' are often used for overlapping packages. A provider should state the log sources, tools, shifts, languages, locations, investigation depth and response actions included in the base service. The managed SOC service should be compared separately from broader managed security services.
The managed SOC provider scorecard
| Criterion | Weight | Evidence to request |
|---|---|---|
| Telemetry and onboarding | 15% | Source inventory, parsing QA, health monitoring and onboarding plan |
| Detection engineering | 20% | ATT&CK coverage, rule lifecycle, test cases and false-positive tuning |
| Investigation quality | 15% | Anonymised case record showing timeline, evidence and analyst reasoning |
| Response authority | 15% | Action matrix for isolate, disable, block, collect and escalate |
| Service reliability | 10% | Staffing, handover, failover, queue monitoring and continuity evidence |
| Measurement and reporting | 10% | Event-to-case funnel, MTTD/MTTR definitions and missed-detection review |
| Data governance | 10% | Data location, access, retention, sub-processors and evidence deletion |
| Exit readiness | 5% | Rule, data, case, dashboard and knowledge transfer on termination |
Weight the capabilities that change incident outcomes
Demand a proof of value, not a dashboard tour
- Provide a representative sample of endpoint, identity, cloud, network and application telemetry with known quality issues.
- Agree five to ten attacker behaviours or incident scenarios that matter to your environment.
- Measure whether the provider ingests the data correctly, detects the behaviour, investigates it and produces a useful decision.
- Review false positives, missed detections, escalation clarity and the time spent at each stage rather than one average SLA.
- Require a remediation plan for visibility gaps before converting the exercise into a long-term contract.
SLA questions that change the answer
- When does the detection clock start: event time, ingestion time, alert creation or analyst assignment?
- What pauses the response clock, and are customer delays excluded from the published metric?
- Does the SLA measure an automated notification or a human-validated investigation?
- Which severities are covered, and who decides severity when facts change during the incident?
- What service credit applies, and what corrective review follows a missed or late detection?
- Which containment actions are pre-authorized and which wait for a named customer decision-maker?
What a useful monthly report contains
| Report section | What it should answer |
|---|---|
| Data health | Which expected sources were missing, delayed, noisy or incorrectly parsed? |
| Detection funnel | How did raw events become alerts, cases, incidents and confirmed threats? |
| Coverage | Which priority techniques are detected, tested, partially covered or blind? |
| Cases | What happened, what evidence supports it, what action was taken and what remains open? |
| Service quality | What were the queue, acknowledgement, investigation and response times by severity? |
| Improvement | Which rules, parsers, playbooks and controls changed because of the month's findings? |
Commercial and technical red flags
- Pricing based only on EPS or data volume without an inventory of sources, retention and investigation workload.
- A large rule count with no test cases, owner, lifecycle or ATT&CK coverage view.
- A 24×7 claim that describes alert receipt but not human investigation and escalation coverage.
- No written response-action matrix, leaving containment authority ambiguous during a live incident.
- Metrics that count every automated alert as a detection and every email as a response.
- No exit plan for rules, parsers, case history, dashboards, threat context and knowledge transfer.
The RFP questions to send every provider
- List every included log source and the process for validating parser and timestamp quality.
- Show how detections are written, tested, tuned, approved, deployed and retired.
- Provide the investigation and response workflow for identity compromise, ransomware and cloud key exposure.
- State analyst coverage, shift handover, surge capacity and continuity arrangements.
- Define MTTD, MTTA, investigation time and MTTR precisely, including every clock pause.
- Describe data residency, privileged access, retention, sub-processors and termination deletion.
- List all base-service exclusions and the rate card or process for work outside those boundaries.
Review the telemetry, detection, investigation, response, governance and handover model before choosing a platform or contract length.
