Blog

min read

Best AI Red Teaming Providers for Enterprise AI Agents and GenAI Apps (2026): A Buyer's Comparison

By

Dor Sarig

and

September 16, 2026

min read

The best AI red teaming providers for enterprise AI agents and GenAI applications in 2026 are those that test the deployed system rather than only the underlying model, repeat the test on every release, and deliver proof of what executed rather than a score. This guide compares eleven providers and tools against the OWASP vendor-evaluation criteria for AI red teaming, states what each is best for, and includes Pillar's Red Graph Suite under the same rubric with its limitations stated.

What is AI red teaming, and what is it not?

AI red teaming is adversarial testing of an AI system, its model, prompts, retrieval, tools, memory and permissions, to find the ways a motivated attacker could make it leak data, take an unauthorized action, or fail its users. It differs from a penetration test in what is under test: a pentest probes code paths and infrastructure, while an AI red team probes behavior that changes with every prompt edit, tool addition, knowledge-base refresh and model update.

The OWASP GenAI Security Project's red teaming manual puts the distinction in one line: prompt injection is not a vulnerability, it is an attack vector. The vulnerabilities it reaches are inadequate permission controls, missing data-access boundaries, insufficient output validation and unvalidated external content. A red team that evaluates only model output tests the vector and misses the vulnerability. The relevant question is what the system did next.

This framing underpins the evaluation criteria below.

How we compared the providers

OWASP published its Vendor Evaluation Criteria for AI Red Teaming Providers and Tooling in January 2026 under a Creative Commons license. It lists thirteen evaluation dimensions, six green flags and six red flags. We consolidated those into seven questions a buyer can answer from a demo and a scoped trial, and scored every provider in this guide, including Pillar, against them using each vendor's published documentation.

#QuestionWhy it mattersOWASP flag it maps to
1Where does it attack from?Model endpoint, API, or the application's real interface?Agentic and RAG risk lives in tools, retrieval and permissions, which are only reachable through the system, not the model.Red flag: "focus exclusively on model outputs"
2Is it multi-turn and stateful?The average successful attack takes five turns; memory and session attacks need state.Green flags: "reproducible single-turn and multi-turn evaluations"; "ability to evaluate stateful systems"
3Does it test for data exposure specifically?90% of successful attacks in production leaked data. Exfiltration through fetched URLs, tool outputs and RAG poisoning are the enterprise-grade cases.Green flag: "custom testing with novel findings"
4Is it continuous and version-aware?An AI system changes weekly; a point-in-time report expires with the next deploy.Green flag: "reproducible evaluations"; red flag: "stock jailbreak libraries"
5What evidence ships with a finding?A score, a transcript, or proof of execution?Findings get argued away without reproduction steps.Green flag: "clear reporting with business impact mapping"; red flag: "black-box scoring"
6Which frameworks, and which version?Compliance teams need OWASP, MITRE ATLAS and NIST AI RMF mapping, and the OWASP LLM Top 10 renumbered in August 2026.Dimension 13: "legal, ethical and compliance posture"
7Does it close into runtime enforcement?A finding that becomes a guardrail policy is fixed; a finding in a PDF is a ticket.Green flag: "actionable remediation guidance"

Comparison at a glance

ProviderTypeAttacks fromMulti-turnContinuousData-exposure testsFrameworks statedRuntime pairingBest for
Pillar Red Graph SuiteAutomated, agenticApplication's real UI, plus discovered endpointsYesVersion-controlled per releaseYes: exfiltration, tool and data-access chains, proof of what leftOWASP Agentic Top 10, OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, SAILSame platform; findings flow into guardrail policyAgentic and RAG applications: threat modeling, continuous red teaming and runtime guardrails in one lifecycle, with data-exposure testing as the core
Lakera Red (Check Point)AutomatedModel or agent API endpointYes, against stateful agent endpointsScan-basedYes: data exfiltration and PII, system-prompt and tool extractionNot statedLakera Guard, via documented manual mappingLLM application testing paired with Guard, with broad safety and brand-risk objectives and 100+ languages
HiddenLayerAutomatedAI applications and modelsNot stated"Continuously test"Not stated"OWASP-aligned"AI Runtime Security (AIDR)Model and supply-chain security programs that want red teaming as one module of a platform
Palo Alto Prisma AIRS AI Red Teaming (Protect AI)AutomatedModels, applications and agents via OpenAI, Bedrock, Databricks, Hugging Face, REST and Copilot Studio connectorsYesScheduled scansYes: system-prompt and tool leak, indirect injectionOWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, DASFAIRS Runtime SecurityPalo Alto shops that want red teaming under the same license and credits as the rest of the stack
Zscaler AI Red Teaming (SplxAI)AutomatedApps, RAG chatbots, agentic workflows and LLM APIs via no-code connectorsYesCI/CD-triggerableYes: data exfiltration and RAG poisoning probes12 frameworks including OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, ISO 42001, EU AI ActZscaler AI Guard and Bedrock Guardrails policy generationCompliance-heavy programs that want the widest framework reporting and automated system-prompt hardening
Noma AI Red TeamAutomated, agenticAny agentic endpoint, including in production with SSOAdaptive sequencesPosture-triggeredYes: RAG exploitation, tool misuse, MCPNot stated on the red-team pageNoma Runtime, as signatures and policiesAISPM-led programs where posture findings should trigger targeted tests
Straiker Ascend AIAutomated, agenticCI/CD, staging and productionYes24/7Named in research; product coverage not itemizedOWASP LLM Top 10, OWASP Agentic Top 10, MITRE ATLAS, NIST AI RMF, EU AI ActDefend AI, attacks converted to guardrail rulesTeams that want paired attack and defense agents running continuously
MindgardAutomatedAPIs, CI/CD, Burp SuiteNot statedYesNot statedNot statedNoneResearch-led continuous testing with attacker-style reconnaissance, in security-team tooling
DreadnodeAutomated platform plus open-source SDKEndpoints via TUI, CLI or Python SDKYesVia CIYes: exfiltration and MCP/tool attack transformsOWASP LLM Top 10, OWASP Agentic, MITRE ATLAS, NIST AI RMF, Google SAIFNoneOffensive security teams that want to extend the attack catalog themselves
SynackManaged, human plus AIApplication and APIHuman-ledContinuous engagementScoped per engagementPer engagementNoneHuman-validated, managed testing where FedRAMP or regulator sign-off is required
Microsoft PyRIT and AI Red Teaming Agent (preview)Open source, plus Azure Foundry serviceModel and Foundry-hosted agent endpointsCrescendo and multi-turn orchestrators; agentic categories single-turnManual or scriptedSensitive-data leakage category (agents, cloud only)None built inNoneAzure AI Foundry teams that want a free starting point

"Not stated" means the vendor's public documentation does not make the claim as of September 2026. It is not a claim that the capability is absent.

The providers

Providers are grouped as OWASP groups them: automated platforms, managed human red teams, and open-source frameworks. Pillar is listed first and scored on the same criteria.

1. Pillar Security Red Graph Suite

Best for: agentic and RAG applications: threat modeling, continuous red teaming and runtime guardrails in one lifecycle, with data-exposure testing as the core.

Pillar's Red Graph Suite starts from the premise in the OWASP manual: test the system, not the model. Autonomous attack agents connect to an AI application through its own interface, by URL or user login, with no code changes, and begin with reconnaissance: system prompt, tools, knowledge sources, permissions. That map is rendered as a graph of agents, tools and data, so a permission that is safe on its own shows up as dangerous once it is chained to a tool that can fetch a URL or run a query. The agents then run multi-turn attacks through the real workflow, adjusting tactics as the target responds, and every finding ships with a video of the attack, a step-by-step log of what the agent did, and the full transcript. In a recent assessment of an AI helpdesk agent, Red Graph Suite achieved 15 of 17 attack objectives through the agent's chat interface, including command execution on the host and data exfiltration through a fetched web page, and re-ran the same objectives against six subsequent releases.

Results are tracked across releases. Each objective's attack success rate is version-controlled, and findings are labeled stable, newly found, hardened or regressed, so the security team watches posture move rather than reading a report that expired at the next deploy. Findings map to the OWASP Agentic AI Top 10 and LLM Top 10, MITRE ATLAS, NIST AI RMF and Pillar's own SAIL framework, co-developed with security leaders from Google, Microsoft, AT&T, SAP, Salesforce, ServiceNow, JPMorgan Chase and others. Because Red Graph Suite is one part of Pillar's platform, findings flow into runtime guardrails, closing the loop from discovering an attack path to enforcing a policy against it. Gartner profiled the approach in its July 2026 report on the coolest vendor innovations in AI software security.

Where it is strongest: data-exposure and tool-abuse testing in agentic and RAG systems, proof of execution, and posture tracking across releases. It supports Microsoft Copilot endpoints, homegrown applications, any agent with a web interface, and endpoints discovered through the Wiz integration.

Limitations: Red Graph Suite is built to attack an application, so it needs an application to log into. Teams testing a bare model endpoint will start faster with the open-source tools below. Its attack catalog is curated rather than crowd-sourced, so the count is smaller than the large public libraries, and it is not a managed human service.

2. Lakera Red (Check Point)

Best for: LLM application adversarial testing paired with Lakera Guard, with unusually broad safety and brand-risk objectives and 100-plus-language coverage.

Lakera Red is the offensive half of Lakera's platform, now part of Check Point, alongside Lakera Guard for runtime. Its architecture is Targets, Scans and Results: a target is a model connection or an agent endpoint, which can be stateless (OpenAI-compatible) or stateful with a session identifier, so multi-turn attacks against a conversational agent are supported. Its documentation lists 23 default attack objectives across three categories: security (instruction override, system-prompt extraction, tool extraction, data exfiltration and PII), safety, and a "responsible" category that covers misinformation, copyright, brand damage, unauthorized discounts and hallucination. That third category is wider than most competitors' and matters to consumer-facing deployments. Each scan produces a risk score from the share of harmful evaluations and exports full conversations as JSON.

Lakera's threat intelligence is its differentiator: tens of millions of attack data points growing by roughly 100,000 a day, fed by Gandalf, its public red-teaming game. Limitations to check: Lakera Red tests at the API endpoint rather than through an application's interface, its public documentation states no framework mapping for Red, and the Red-to-Guard integration is a documented manual process in which a team reviews findings and then configures Guard policies by severity.

3. HiddenLayer

Best for: model and supply-chain security programs that want red teaming as one module of a broader platform, including federal environments.

HiddenLayer's platform spans AI discovery, supply-chain security, attack simulation and runtime security, and its red teaming is described as continuously testing agentic and generative AI applications "with adversarial simulations to uncover vulnerabilities before attackers do." Red teaming is marketed as OWASP-aligned and sits next to the Model Scanner, which checks model files for malware and unsafe serialization, and AI Detection and Response for runtime. The company reports more than 50 CVEs disclosed through its research, federal traction including a Missile Defense Agency contract, and a recent $100 million Series B. In August 2026 it added Agent Harness Security for AI coding agents.

Limitations to check: HiddenLayer's public materials do not itemize the red-teaming module's attack interface, multi-turn depth, data-exposure coverage or framework mapping beyond "OWASP-aligned." Buyers who need those specifics should ask for them in the evaluation; the platform's center of gravity is model and supply-chain protection.

4. Palo Alto Networks Prisma AIRS AI Red Teaming (Protect AI)

Best for: Palo Alto customers who want red teaming under the same license and credit model as the rest of their stack, with the largest published attack library.

Prisma AIRS AI Red Teaming, which originated in the Protect AI acquisition and requires its own license, targets models, applications and agents through connectors for OpenAI, Hugging Face, Databricks, AWS Bedrock, REST and Microsoft Copilot Studio, with a lightweight network channel to reach private endpoints. It runs three scan types: an attack library refreshed every two weeks from Unit 42 and the huntr community of more than 18,000 researchers, with over 500 vectors across more than 50 techniques; an LLM-driven agent scan that adapts to the target and can run fully automated or human-augmented; and custom prompt sets. Categories include indirect prompt injection, multi-turn, system-prompt and tool leakage, remote code execution and evasion encodings, and reports map to the OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF and Databricks' DASF. Palo Alto claims setup in under ten minutes and reports within five hours.

Limitations to check: continuous testing is available as scheduling rather than a documented continuous mode, and the product is priced through Software NGFW credits with API usage billed by token, which is transparent but ties the red team to the firewall budget. Documentation notes that AI traffic is sent to a US region for inspection, which matters for data-residency reviews.

5. Zscaler AI Red Teaming (SplxAI)

Best for: compliance-heavy programs that want the widest framework reporting and automated system-prompt hardening, especially existing Zscaler customers.

Built on the SplxAI acquisition, Zscaler AI Red Teaming targets AI applications, RAG chatbots, agentic workflows and LLM APIs through no-code connectors including Copilot Studio, Salesforce Agentforce, Glean, Slack and Teams, and can be triggered from CI/CD. It ships more than 25 predefined probes across security (including data exfiltration, RAG poisoning, context leakage and code execution), safety, hallucination and business alignment, with multi-turn strategies such as multishot and delayed attacks, and multimodal inputs. Its reporting maps to twelve frameworks, including OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, ISO/IEC 42001 and the EU AI Act, and its remediation features are unusual: automated system-prompt hardening with a line-by-line diff and a policy generator that emits rules for Zscaler AI Guard and AWS Bedrock Guardrails. Zscaler's brochure claims 95% less testing effort and 97% lower cost than manual red teaming.

Limitations to check: the platform attacks through connectors and APIs rather than an application's interface, and its documentation still references the 2025 edition of the OWASP LLM Top 10. Ask which version a compliance report will cite.

6. Noma Security AI Red Team

Best for: AI security posture management programs that want posture findings to trigger targeted red-team tests automatically.

Noma describes its red team as "itself an intelligent agent, building and adapting attacks based on the specific behavior of each application under test," rather than a static library. It claims to test any agentic endpoint regardless of framework, including in production with OAuth, enterprise SSO and custom auth flows; names RAG exploitation, memory manipulation, tool misuse and MCP vulnerabilities as coverage powered by Noma Labs; and closes the loop in two directions, with posture scanners triggering assessments and findings flowing into Noma's runtime protection as detection signatures and guardrail policies.

Limitations to check: Noma's red-team page states no framework mapping (its compliance module is separate), does not explicitly describe multi-turn depth, and does not publish pricing. Independent detail on the product is thinner than for the incumbents above, so a scoped trial matters more.

7. Straiker Ascend AI

Best for: teams that want paired attack and defense agents running continuously across CI/CD, staging and production.

Straiker's Ascend AI is "continuous adversarial agent testing": attack agents run around the clock against agentic applications, with coverage claimed against the OWASP LLM Top 10, OWASP Agentic AI Top 10, MITRE ATLAS, NIST AI RMF and the EU AI Act, and a sibling product, Defend AI, converts successful attacks into guardrail rules. The company, founded by former Prisma Cloud leadership, closed a $64 million Series A in June 2026 and publishes research through STAR Labs; its July 2026 report claimed more than 1,700 successful exploits, with 36% of successful attacks on coding agents reaching remote code execution on the developer machine.

Limitations to check: STAR Labs figures are vendor-published without a released methodology or dataset, and the public materials do not specify whether Ascend attacks through an application's interface or its API. Both should be confirmed during evaluation.

8. Mindgard

Best for: research-led continuous testing with attacker-style reconnaissance, delivered inside the tooling security teams already use.

Mindgard, a spin-out of more than a decade of AI security research at Lancaster University with offices in Boston and London, positions its platform as "an autonomous red teamer" that "uses attacker-style reconnaissance and intelligence powered by its vulnerability disclosures to reveal how adversaries discover and exploit AI agents and systems." It deploys through CI/CD, Burp Suite or a single click, covers agentic workflows and multimodal systems, and cites more than 150 publicly disclosed AI vulnerabilities and tenfold faster assessments.

Limitations to check: Mindgard's homepage does not state framework mappings, multi-turn depth or specific data-exposure coverage, and it has no runtime enforcement product, so findings are remediated elsewhere.

9. Dreadnode

Best for: offensive security teams that want to read, run and extend the attack catalog themselves.

Dreadnode pairs an open-source Python SDK with a commercial platform for analytics, compliance reporting and evidence export, and can be driven from an agentic terminal interface, a CLI for CI/CD, or code. Its May 2026 paper, co-authored by Counterfit co-creator Will Pearce, documents more than 45 attack strategies, including Crescendo, TAP and PAIR, over 450 prompt transforms across 38 modules, including MCP and tool attacks, multi-agent exploits and exfiltration, and more than 130 detection scorers, auto-tagged to the OWASP LLM Top 10, the OWASP Agentic Security Initiative, MITRE ATLAS, NIST AI RMF and Google SAIF. It publishes its benchmarks and datasets, including a July 2026 study in which professional penetration testers scored 4,897 autonomous-agent tool calls for scope.

Limitations to check: Dreadnode has not published a demonstration of an autonomous attacker navigating a deployed agent application end to end through its interface, and it has no runtime product; it is a red team's toolkit, not a closed loop.

10. Synack

Best for: human-validated, managed testing where a regulator, auditor or federal program needs a named tester behind every finding.

Synack is the managed option on this list. Its platform combines Sara, the Synack Autonomous Red Agent that "identifies, validates, and prioritizes vulnerabilities," with the Synack Red Team of more than 1,500 vetted researchers, and offers AI and LLM pentesting alongside application, API and cloud testing on a FedRAMP-authorized platform. Synack's own comparison of the category argues that automated scanners miss novel exploits that human testers catch, and its model is built on that premise.

Limitations to check: managed testing is scoped per engagement, so continuity, data-exposure depth and framework mapping depend on the statement of work rather than a product spec, and cost scales with researcher time. Many enterprises pair a managed engagement with an automated platform rather than choosing between them.

11. Microsoft PyRIT and the Azure AI Foundry AI Red Teaming Agent

Best for: Azure AI Foundry teams that want a free, well-documented starting point before buying a platform.

PyRIT is Microsoft's open-source Python framework for orchestrating adversarial probes against LLM endpoints, with a large set of prompt converters (Base64, ROT13, leetspeak, ASCII art, Unicode confusables and more) and multi-turn orchestrators including Crescendo. The AI Red Teaming Agent, in preview inside Azure AI Foundry, wraps PyRIT with Foundry's risk and safety evaluators and produces a JSON scorecard of attack success rate by category and technique. It covers content risks, code vulnerabilities and indirect prompt injection for models, and adds prohibited actions, sensitive-data leakage and task adherence for Foundry-hosted agents.

Limitations, in Microsoft's own words: the agentic categories are single-turn and English-only, use mock tools rather than real ones, and support Foundry-hosted agents only; workflow agents, non-Foundry agents and browser or computer-use tools are out of scope. Microsoft states that "the adversarial nature of our red teaming evaluations is controlled to avoid real world impact." There is no framework mapping out of the box. NVIDIA's Garak and Promptfoo are the other open-source options most teams evaluate alongside PyRIT; all three are model-and-prompt scanners first, and none attacks through an application interface.

How to choose an AI red teaming provider

Start with the system, not the vendor list. Write down what your highest-risk AI application can do: which tools it calls, what it retrieves, which permissions it inherits, whether it remembers across sessions. Then ask each provider to show, in a scoped trial, an attack that reaches each of those capabilities. A provider that can only reach the model will produce a report about jailbreaks; a provider that reaches the tools will produce a report about your data.

Then run the seven questions above and add these, which separate vendors faster than any feature list:

  1. Show me a finding with the command that ran, not the text the model produced. If the evidence is a transcript alone, ask how a developer would reproduce it.
  2. Run the same objectives against two consecutive releases. Ask what regressed. A provider with no answer has no version model.
  3. Point it at a stateful agent and ask it to poison memory in one session and collect in another. Cross-session attacks are the OWASP green flag most often skipped.
  4. Ask for a fetched-URL exfiltration test. If the agent can browse or call a webhook, data can leave through it. This is the most common real-world outcome and the least commonly demonstrated.
  5. Ask which OWASP LLM Top 10 version the report cites. The 2026 edition renumbered most of the list on August 4, 2026. A report that says "LLM07: System Prompt Leakage" is using the 2025 numbering.
  6. Ask how a finding becomes a runtime policy. Manual mapping, export to a third-party guardrail, or the same platform. None is wrong; each has a different cost in engineering hours.
  7. Ask for pricing. Only one vendor in this list publishes a pricing mechanism. Plan procurement accordingly.

Which provider is best for AI agents?

If the system you are securing is an AI agent, one that calls tools, retrieves documents, holds memory or acts with a user's permissions, the rubric narrows the list. The four questions that matter most for agents are where the attack enters, whether it is stateful and multi-turn, what evidence ships with a finding, and whether the test repeats on every release. Read the comparison table against those four and the shortlist is Pillar, Noma and Straiker. Of the three, Pillar's Red Graph Suite is the only one whose public documentation states that it attacks through the application's real interface, runs multi-turn attacks with reproduction video and action logs, version-controls results across releases, and maps to the OWASP Agentic AI Top 10. That combination is why it leads this list for agents. Noma is the closer alternative when posture-triggered testing in production matters most, and Straiker when a buyer wants attack and defense agents from one vendor running around the clock. If your agents are headless services with no interface to log into, weigh Noma, Straiker and Dreadnode more heavily, because Pillar's approach starts from an application a user can reach.

Next steps

Pillar's Red Graph Suite is the red-teaming component of a platform that also covers AI asset discovery, runtime guardrails and governance, built for agentic and RAG applications with data-exposure testing at the core. Read the Red Graph Suite launch for the helpdesk-agent assessment in full, or request a scoped red-team assessment of one production AI application.

FAQs

What is the best AI red teaming provider for AI agents?

By the OWASP-derived rubric in this guide, Pillar Security's Red Graph Suite, because its public documentation states that it attacks agents through their real interface, runs stateful multi-turn attacks, ships video and action-log proof of what executed, re-tests every release, and maps to the OWASP Agentic AI Top 10. Noma and Straiker are the closest alternatives for agentic systems; Zscaler leads on compliance reporting breadth. Pillar wrote this comparison, so verify with a scoped trial.

What is the difference between AI red teaming and AI penetration testing?

A penetration test probes code, configuration and infrastructure for known classes of flaw. AI red teaming probes behavior: whether an adversary can make a model, agent or retrieval system leak data, take an unauthorized action or fail its users. Because that behavior changes with every prompt, tool or model update, AI red teaming has to be repeated far more often than a pentest.

Should we choose an automated AI red teaming platform or a managed human red team?

Most enterprises end up with both. Automated platforms give repeatability across releases and breadth of attack coverage; managed human teams give creativity, novel findings and a named tester for regulators. OWASP's guidance treats them as different vendor types with different evaluation criteria rather than substitutes.

What should an AI red team test in a RAG application?

Whether poisoned or untrusted documents in the retrieval corpus can change the system's behavior, whether retrieval can be steered to surface data the user should not see, whether fetched content can carry data out, and whether the application's permissions let retrieval reach systems beyond its remit. Testing the model alone finds none of these.

How often should AI systems be red teamed?

On every release that changes a prompt, tool, knowledge source or model, and continuously for systems that change without a release, such as agents with live retrieval or memory. Pillar's State of Attacks on GenAI found the average successful attack took 42 seconds.

Which frameworks should AI red teaming findings map to?

For most enterprises: the OWASP Top 10 for LLM Applications and the OWASP Agentic AI Top 10 for vulnerability classes, MITRE ATLAS for adversary techniques, and NIST AI RMF for governance. Check which OWASP LLM Top 10 version a provider cites, since the 2026 edition renumbered the list.

Why does data exposure matter more than jailbreaks?

Because it is the outcome attackers are after. In production telemetry, 90% of successful attacks resulted in sensitive data leaving the system, and the most common vector was extracting the system prompt or business data rather than producing harmful content. Jailbreaks are often the first step; exfiltration is the goal.

Does AI red teaming replace runtime guardrails?

No. Red teaming finds the attack paths; guardrails stop them in production. The efficient pattern is a loop: red-team findings become runtime policies, and runtime detections become new red-team objectives. Providers differ in whether that loop is manual, exported or native to one platform.

Is Pillar's Red Graph Suite suitable if we only want to test a model endpoint?

It can, but it is designed to attack an application through its real interface and prove what the system did. Teams testing a bare model in isolation will start faster with an open-source framework and move to a platform once the model has tools.

Subscribe and get the latest security updates

Back to blog

MAYBE YOU WILL FIND THIS INTERSTING AS WELL

Pillar Security Named an AI Security Technical Innovator in the Latio 2026 AI Security Market Report

By

Dor Sarig & Ziv Karliner

and

September 17, 2026

News
The Audit Window: What the EU AI Act's Deferral Actually Bought You

By

Dor Sarig

and

September 9, 2026

Blog
Valid, But Never Issued: Session Spoofing and SSRF in Grafana MCP

By

Ariel Fogel

and

September 2, 2026

Research