Can “AI Security” Really Be Tested? The AV-Comparatives Perspective

Artificial intelligence has rapidly become one of the most prominent terms in cybersecurity. Security vendors increasingly advertise AI-powered detection, AI assistants, autonomous security agents and products designed to protect AI systems themselves. Inevitably, this has also created demand for independent “AI security tests”.

At AV-Comparatives, we believe independent testing must follow technological developments. However, a new technology or marketing category does not automatically mean that a scientifically sound test can be built around it.

Before introducing any test, one question has to be answered first:

What exactly are we testing, and can the results be measured objectively, reproducibly and fairly?

At present, we believe that broad “AI security testing” faces substantial methodological limitations, and here is why.

AI is not a single security function

“AI” is an extremely broad term.

AI and machine-learning technologies have already been part of cybersecurity products for years. They contribute to malware detection, behavioural analysis, phishing protection, anomaly detection, EDR/XDR capabilities and many other security functions.

AV-Comparatives does not need a separate “AI test” to evaluate these technologies, because we test the outcome.

If a vendor claims AI improves malware detection, the relevant question is whether the product detects malware while maintaining an acceptable false-positive rate. If AI improves endpoint prevention, the relevant question is whether attacks are prevented. If AI assists EDR capabilities, the relevant question is whether malicious activity is detected and whether useful telemetry is generated.

Whether the underlying technology relies on machine learning, neural networks, rules, signatures, behavioural models, reputation systems or a combination of all of these is secondary to the protection ultimately delivered to the user.

In most modern security products, these technologies work together as part of one integrated protection stack. From an external testing perspective, it is generally not possible to determine whether a particular detection or prevention was the result of an AI component, another protection mechanism, or a combination of several technologies.

AV-Comparatives therefore evaluates the security outcome of the product as a whole. If AI contributes to that outcome, its contribution is inherently reflected in the product’s performance, but it cannot normally be isolated or attributed specifically to AI.

Consequently, AI-related capabilities are already considered in a number of AV-Comparatives tests and reviews wherever they form part of the product or functionality being evaluated.

AI is therefore not necessarily something that requires a separate test. It is often one of many technologies contributing to a security outcome that can already be tested.

Testing AI itself is a different problem

The situation becomes more complicated when the subject is an AI model or an agentic system itself.

The result of an interaction can depend on several variables at once, including:

  • The model and model version
  • The system prompt
  • The agent framework and available tools
  • Permissions and memory
  • Context and external information
  • Configuration

Unlike many traditional security mechanisms, AI and agentic systems can produce different decisions, actions or outcomes from identical inputs.

Models can also change rapidly. Cloud-based systems may be updated without the tester controlling, or even knowing, every detail of the change.

Consequently, a result obtained today with a particular model, agent and configuration may not necessarily be reproducible several weeks later, or transferable to another environment.

The fundamental question therefore remains: what exactly would an “AI security test” measure?

Without a narrowly defined security claim and a clearly measurable outcome, the term itself says very little.

Non-determinism creates a fundamental testing problem

Traditional security testing also contains variability, but AI systems introduce another level of non-determinism.

The same attack may succeed during one execution and fail during another. The underlying model itself may refuse an instruction before a security product ever has the opportunity to intervene.

A tester therefore has to determine whether an attack was prevented by the security product, refused by the model, affected by the agent environment, or simply the result of a different model decision.

Repeated executions and statistical analysis can reduce this uncertainty, but they do not eliminate the underlying attribution problem.

A percentage score can look precise while representing a highly environment-dependent result.

Precision of presentation should not be confused with scientific certainty.

An “AI test” can quickly become obsolete

Another challenge lies in selecting representative test cases.

Agentic systems can be attacked through several mechanisms, among them:

  • Direct prompts and indirect prompt injection
  • Malicious retrieved content
  • Tools, plugins and MCP servers
  • Memory manipulation
  • Cross-agent communication

These attacks can also depend heavily on the specific model and agent architecture in use.

A test built around a fixed collection of prompts may therefore primarily measure how one particular system handles those particular prompts in that particular environment.

The field is also developing so rapidly that attack techniques, models and defensive mechanisms can shift significantly within a short period of time.

A test environment or set of scenarios that looks representative today may therefore offer only limited insight into a different model, a different agent architecture, or a future version of either.

Emerging guidelines are valuable, but not yet a mature testing methodology

Emerging industry guidance for testing AI and agentic security is important and welcome. This work helps identify relevant attack vectors, terminology, testing environments and reporting requirements.

It also recognises several of the challenges described above, including model dependency, agent-host dependency, non-determinism and the limited maturity of some agentic attack scenarios.

These efforts are useful steps towards future testing methods. However, defining general principles for how testing should be approached is not the same as having a mature methodology that produces reproducible, representative and meaningful security results.

AV-Comparatives believes that this distinction matters.

What can meaningfully be tested today?

None of this means that security functionality related to AI cannot be tested.

The key is to test a specific security claim, rather than “AI” as an abstract category. Clearly defined functional and scenario-based assessments can already provide useful information.

For example, a product claiming to protect an AI agent against indirect prompt injection could be exposed to controlled malicious content, with the assessment determining whether the attack results in an unauthorized action or data disclosure.

Products could similarly be evaluated against defined scenarios involving malicious tool usage, prompt injection, unauthorized data access or attempted data exfiltration.

Such assessments can answer a focused question:

Did the security control prevent this defined attack under these defined conditions?

That is a meaningful and measurable result.

What such a test does not establish is that “AI security” as a whole has been tested.

For this reason, AV-Comparatives currently considers focused functional assessments, individual product reviews and dedicated research projects more appropriate for emerging AI-security technologies than broad “AI security tests”.

Outcome matters more than the AI label

“AI-powered” has become a widely used marketing term. The presence of AI in a product does not by itself demonstrate better security, and attaching the word “AI” to a test does not automatically create a new testing discipline.

Independent testing should avoid validating terminology and instead validate outcomes.

If AI genuinely improves a security product, that improvement should ultimately become visible through better protection, better detection, fewer false positives, improved response capabilities or other measurable benefits. Those benefits can, and should, be tested. The underlying technology used to achieve them does not necessarily require a separate “AI test”.

Our approach to AI security testing

AV-Comparatives continuously evaluates new technologies and attack techniques, and incorporates relevant developments into its methodologies wherever appropriate.

We are also following developments around AI agents, prompt injection, agentic security controls and emerging testing frameworks closely.

However, we believe independent testing has a responsibility not to create tests simply because a subject is receiving significant market attention.

A test should have a clearly defined subject, measurable security outcomes and a methodology that is sufficiently objective, repeatable, representative and fair.

As AI technologies and the corresponding testing science mature, AV-Comparatives will continue to evaluate and develop new methodologies wherever they can provide meaningful, reproducible and useful security insights.

Independent testing should measure what can be demonstrated, not what is currently fashionable.

Andreas Clementi, CEO & Co-Founder, AV-Comparatives