An AI security assessment should test the system that can disclose data or take action, not only the model that produces text. The provider must follow hostile input through instructions, retrieval, memory, tools, identities, approvals and downstream APIs; reproduce the impact; and leave regression tests your team can rerun after a model, prompt, connector or guardrail changes.
Define the AI system before comparing providers
| Layer | Inventory before scoping | Failure to test |
|---|---|---|
| Model and instructions | Models, system prompts, routing, policies and fallback paths | Jailbreak, goal hijack or unsafe output |
| RAG and knowledge | Sources, ingestion, chunking, retrieval, tenancy and citations | Poisoning, cross-tenant disclosure or hidden instructions |
| Tools and agents | Functions, MCP servers, APIs, code execution and action limits | Unauthorized side effects or command injection |
| Identity | Users, tenants, service accounts, delegation and approval roles | Confused deputy or privilege escalation |
| Memory | Conversation, profile, task and long-term stores | Persistent injection or another user's data exposure |
| Operations | Logging, budgets, rate limits, shutdown, review and incident response | Undetected abuse, runaway cost or unsafe recovery |
Use the agentic AI security testing checklist to create the inventory and the MCP server security guide when tools use MCP. The commercial AI and LLM penetration testing service should then map every in-scope trust boundary to an attack method and evidence output.
AI security provider scorecard
| Criterion | Weight | Evidence to request |
|---|---|---|
| System coverage | 20% | Architecture-led plan spanning model, data, tools, identity, memory and operations |
| Application security depth | 15% | API, authorization, injection, SSRF, secrets and business-logic testing |
| AI attack depth | 20% | Direct/indirect injection, RAG poisoning, tool misuse and multi-step abuse cases |
| Impact validation | 15% | Reproducible paths to disclosure, unauthorized action, persistence or cost |
| Safety and data handling | 10% | Test boundaries, sensitive prompts, output controls and evidence governance |
| Remediation quality | 10% | Control placement at identity, tool, data and application layers |
| Regression handover | 10% | Versioned test cases, expected outcomes and CI-ready rerun guidance |
Weight end-to-end risk more heavily than prompt volume
What a complete assessment should test
- Direct and indirect prompt injection across user input, retrieved documents, web pages, tickets, emails and tool output.
- System-prompt leakage, goal hijacking, instruction conflicts and attempts to disable or reinterpret policy.
- Tenant and object authorization at every data source and tool, independent of what the model believes is permitted.
- RAG poisoning, retrieval manipulation, citation spoofing, embedding or metadata leakage and cross-tenant retrieval.
- Tool misuse, unsafe parameters, injection into downstream APIs, server-side request forgery and excessive agency.
- Memory poisoning, persistence across sessions, shared-context exposure and unsafe deletion or retention behaviour.
- Model, prompt, plugin, connector and package supply-chain changes that alter the trusted behaviour of the system.
- Rate limits, budget exhaustion, output handling, logging, human approval, pause, revoke and incident-recovery controls.
Ask for evidence, not a jailbreak count
| Weak evidence | Decision-useful evidence |
|---|---|
| The model answered a prohibited question | The output, affected policy, repeatability, severity and realistic user impact |
| A prompt injection succeeded | The untrusted source, instruction path, changed decision and downstream consequence |
| The agent called a tool | The identity used, authorization check, arguments, side effect and audit trail |
| RAG leaked data | The tenant boundary crossed, retrieval path, exposed fields and root control failure |
| A guardrail blocked the test | Whether the underlying tool, data and authorization controls still prevent the action |
| Hundreds of prompts were run | Coverage by abuse case, trust boundary, expected result and regression status |
Questions for an AI security provider
- How will you test authorization at the data and tool layers rather than asking the model whether an action is allowed?
- Which indirect-injection sources will you place hostile content into, and how will you trace influence through the plan?
- How will you test cross-tenant access without retaining or reproducing real customer data?
- Which parts are automated, which require manual multi-step analysis and how are false positives reviewed?
- How will you validate tool misuse safely when a function can send, publish, purchase, deploy, modify or delete?
- What changes trigger a retest: model version, prompt, retrieval corpus, tool schema, permission, memory or guardrail?
- What regression artifacts will our engineering team receive, and can they run them in CI without your platform?
Common red flags
- The scope is described only as jailbreaking the model or checking one AI risk list.
- The provider does not ask for data flows, tool permissions, tenant model, identities or downstream actions.
- A proprietary scanner score replaces reproducible findings and business-impact validation.
- The remediation advice is limited to changing the system prompt or adding a generic output filter.
- No application-security testing is included for APIs, authorization, secrets, injection and business logic.
- No versioned regression suite is delivered even though the model and surrounding system change frequently.
A fair proof-of-capability exercise
- Provide a small non-production deployment that includes retrieval, one sensitive tool, at least two roles and realistic approval gates.
- Define three abuse outcomes, such as cross-tenant retrieval, unauthorized tool use and persistent instruction poisoning.
- Require the provider to show the root control failure, not merely a surprising model response.
- Compare the clarity, repeatability, safety and remediation value of the findings using the same evaluation rubric.
- Ask how each test becomes a regression check after the fix and what future system changes invalidate the result.
Review an assessment method that covers prompts, RAG, tools, identity, APIs, memory, approvals and downstream impact with reproducible findings and a remediation retest.
