Blog

16

min read

On-Premise AI Security in 2026: How to Secure AI Without Handing Your Data to Another Vendor

By

Dor Sarig

and

October 5, 2026

16

min read

AI models, tools, data and agents protected inside a locked perimeter in the customer's own environment, with the connection to an outside cloud blocked

On-premise AI security means the tool that inspects your prompts, model responses and agent tool calls runs inside your own infrastructure. Your sensitive data is never sent to the security vendor or to a third-party model. For banks, insurers and healthcare providers, that takes a sub-processor, a cross-border data transfer and a third-party risk review out of every AI project. Encrypting traffic to a SaaS tool doesn't achieve the same thing.

Our recommendation. If your AI traffic carries regulated data, such as account numbers, claims, patient records or source code, require three things from any AI security platform. Inspection runs on your own compute. Detection doesn't call an external model. Logs and findings land in a database you own. Pillar's self-hosted deployment, Pillar Stack, does all three inside your Kubernetes cluster. It covers runtime guardrails, AI discovery, red teaming and agentic endpoint posture in one deployment.

Key takeaways

  • Your AI security tool sees more sensitive data than almost anything else you run. To catch a prompt injection or a PII leak, it has to read the full prompt, the full response and every tool call. A SaaS guardrail therefore receives a copy of your most sensitive AI traffic.
  • Encryption doesn't change who is processing the data. US health regulators say a cloud vendor that stores ePHI is a HIPAA business associate even if it can't read the data and holds no key. GDPR, DORA, NYDFS Part 500 and APRA CPS 230 all treat that vendor as a third party you must assess, contract and monitor.
  • Self-hosting removes the vendor from the data flow, not your obligations. You still need access controls, retention rules and evidence. The difference is that they run on your systems, under your policies, without another company in the chain.
  • "On-prem" means different things to different vendors. Ask what runs locally, what still needs outbound access, which detectors call an external model, and which features are cloud-only. The checklist below has the ten questions.
  • Self-hosting has real costs. You run GPU nodes, you schedule upgrades, and some features arrive in the cloud first. Pick self-hosted when you have a hard data requirement, and hybrid when only some workloads do.

What is on-premise AI security?

On-premise AI security is a deployment model in which an organization's AI security controls run on infrastructure it controls rather than in the vendor's cloud. These are the services that discover AI assets, test them, and inspect and enforce policy on live AI traffic. The inspection engine, the detection models, the logs and the findings all stay inside the organization's boundary.

The terms get used loosely. This is how we use them:

Deployment modelWhere inspection runsWhere logs and findings liveWho else sees your AI traffic
SaaS (managed cloud)Vendor's cloudVendor's cloudThe vendor and its sub-processors
Self-hosted in your cloud (VPC)Your cloud account, in your Kubernetes clusterYour database, in your accountNo one outside your organization; your cloud provider as today
On-premise (data center)Your own hardwareYour own storageNo one outside your organization
Air-gappedYour hardware, with no outbound internetYour storageNo one; updates are imported manually
HybridSelf-hosted for regulated workloads, SaaS for the restSplit by workloadDepends on the workload

Most enterprises that say "on-prem" in 2026 mean the second row: the vendor's software running in their own AWS, Azure or Google Cloud account, next to the AI applications it protects. That keeps data under the same contracts, keys and residency controls the organization already has with its cloud provider.

An in-region SaaS solves residency, not access. A vendor that opens a data center in your country keeps your data in-country, but the vendor still receives it, operates it and can be compelled to produce it. For many banks, insurers and health systems, residency is the legal minimum. The real requirement is that no one outside the organization handles the data.

Why does it matter where your AI security tool runs?

Because the security layer is the one component that has to read everything. A guardrail that can't see the prompt can't detect an injection. A data-loss control that can't see the response can't catch a leaked account number. A red-teaming engine that can't reach your internal app can't test it.

Here is what an AI security platform typically handles:

ComponentWhat it readsWhy that is sensitive
Runtime guardrailsEvery prompt, response, retrieved document, tool call and tool resultCustomer PII, PHI, account and card data, internal documents
Session logs and audit trailFull conversations, verdicts, user identitiesA searchable record of everything employees and customers asked your AI
AI discoverySource code, prompts and model references in your repositoriesIntellectual property, system prompts, secrets
Red teamingYour application's behavior under attack, including what data it leaksA map of your weaknesses
Endpoint postureAI agent configurations on developer machinesInternal tool names, MCP servers, file paths

When that platform is SaaS, every row above becomes data you send to another company. The vendor is now a processor under GDPR, likely a business associate under HIPAA, a third-party service provider under NYDFS Part 500, an ICT third-party service provider under DORA and a material service provider candidate under APRA CPS 230. Each label brings contracts, assessments, registers and audits. The aim was to reduce AI risk, and the security tool has added a vendor to the compliance chain.

Which regulations make third-party AI data processing harder?

None of these rules says "you must run AI security on-premise". What they do is make every vendor that touches regulated data a party you have to justify, contract, monitor and be able to exit. Self-hosting changes which vendors are in scope.

RegulationWho it applies toWhat it requires for third partiesWhat self-hosting changes
GDPR (EU, UK GDPR)Anyone processing EU or UK personal dataA processor contract under Article 28, sub-processor approval, and transfer safeguards under Chapter V if data leaves the EEAThe AI security vendor doesn't receive the personal data in your prompts, and no transfer takes place
HIPAA (US)Covered entities and their business associatesA business associate agreement with any vendor that creates, receives, maintains or transmits PHI. HHS says this applies even when the vendor cannot decrypt the dataA vendor that never receives PHI from your AI traffic isn't in that flow (confirm with counsel)
GLBA Safeguards Rule (US)Financial institutionsSelecting and overseeing service providers that can access customer informationCustomer data in prompts stays in your environment
PCI DSS v4.0Anyone storing, processing or transmitting cardholder dataManaging third-party service providers that touch cardholder data, or could affect its security (Requirement 12.8)Card data in prompts isn't sent to the guardrail vendor
DORA (EU)EU banks, insurers, investment firms and other financial entities, since January 17, 2025ICT third-party risk management, contract terms, exit strategies, and an annual register of information on every ICT providerThe vendor supplies software you run; it doesn't operate a service holding your data (it may still belong in the register as a software supplier)
NYDFS 23 NYCRR 500.11New York-regulated banks, insurers and financial services firmsWritten policies for third-party providers that hold or can access nonpublic information: due diligence, MFA, encryption and breach noticeNonpublic information in AI traffic isn't held by the vendor
APRA CPS 230 (Australia)Australian banks, insurers and superannuation trustees, since July 1, 2025Managing material service providers, including a register filed with APRA and fourth-party riskData residency and fourth-party chains stay under your control
EU AI ActProviders and deployers of AI systems in the EULogging, record-keeping and documentation. Article 50 transparency duties have applied since August 2, 2026; Annex III high-risk obligations now apply from December 2, 2027Logs and evidence sit in your systems, under your retention policy (what the deferral actually bought you)

This table summarizes public regulatory text as of October 2026. It isn't legal advice; your counsel decides how each rule applies to your architecture.

On-premise AI security for banks, insurers and healthcare

Banks and financial services

Banks are putting LLMs into fraud operations, customer service, KYC review, internal search and developer tooling. Prompts carry account numbers, transaction histories, customer correspondence and source code. DORA, NYDFS Part 500, GLBA and APRA CPS 230 all ask the same question of every new vendor: what data do you hold, where, and how do we leave? A self-hosted guardrail turns that into a software procurement rather than a data-processing relationship. It also keeps the audit trail, the record of every flagged prompt and blocked response, inside the bank's own retention and e-discovery systems.

Insurance

Insurers use AI on claims notes, medical reports attached to claims, underwriting files and policyholder chat. A single claim file can hold health data, financial data and personal data at once. That means GDPR, HIPAA for health-adjacent lines, state insurance data security laws and, in the EU, DORA. Running inspection in your own environment means the claim text a guardrail reads never becomes another vendor's record.

Healthcare and life sciences

Clinical documentation assistants, patient messaging, prior-authorization agents and research copilots all handle PHI. Under HIPAA, any vendor that receives that PHI needs a business associate agreement, and encryption doesn't change that. Self-hosted guardrails let a health system mask PHI before it reaches a model, block unsafe outputs, and keep a full audit trail of AI interactions, without adding a business associate for the security layer.

Is self-hosted AI security as good as SaaS?

It can be, but it isn't free. Be clear about the trade-offs before you choose:

  • You run the infrastructure. Real-time AI detection runs on GPUs. A production self-hosted deployment needs GPU-enabled nodes, capacity planning and an on-call owner.
  • You own upgrades and patching. New detectors and fixes arrive as releases you apply. Ask how often releases ship, how they're delivered (Helm, GitOps), and how the vendor reports vulnerabilities in its own images.
  • Some features arrive in the cloud first. Capabilities that depend on managed model serving or cross-customer learning may be cloud-only or later on self-hosted. Ask for the feature matrix in writing.
  • Threat intelligence still has to reach you. A self-hosted engine is only as current as its last update. Find out what the vendor ships with each release, and whether anything flows back.
  • Latency usually improves. Inspection runs next to your applications and models, with no round trip to an external API.

When to choose which: choose SaaS unless you have a hard requirement. Choose self-hosted when regulated data flows through your AI applications, when data residency is mandated, or when your third-party risk process makes a new data processor a multi-month project. Choose hybrid when only some workloads carry that data.

What should a self-hosted AI security platform include?

Twelve requirements to put in an RFP:

  1. Local inspection. Prompt, response and tool-call scanning runs on your compute.
  2. Local detection models. Core detectors (prompt injection, jailbreak, PII, PCI, secrets, toxicity) run as models inside your cluster. Any detector that relies on a large language model can be pointed at a model endpoint in your own cloud account or switched off, and never defaults to the vendor's API.
  3. Your data stores. Sessions, findings and audit logs go to databases and queues you operate, with your encryption keys and backups.
  4. A stateless option. A mode that inspects and returns a verdict without persisting any content, for the most sensitive traffic.
  5. A documented egress list. Every outbound domain the deployment needs, what goes there, and whether prompt content is ever included.
  6. Opt-in telemetry. Nothing is sent to the vendor unless you turn it on.
  7. Your identity and secrets. SSO, SCIM, workload identity (IRSA or Pod Identity on AWS) and native Kubernetes secrets, with no hard dependency on one secrets manager.
  8. Coverage beyond the gateway. Discovery of AI assets in code, red teaming of internal applications, and AI agent posture on developer machines. A runtime filter alone leaves most of the attack surface unseen.
  9. Integration with your stack. SIEM export, ticketing and your AI gateway (LiteLLM, Kong, TrueFoundry, Agent Router) without routing traffic through the vendor.
  10. GitOps-friendly delivery. Helm charts, Flux or Argo CD support, and versioned upgrades.
  11. High availability. Multi-replica services and external, clustered Postgres and Redis for production.
  12. A written feature matrix. What works self-hosted today, what's cloud-only, and what's on the roadmap with dates.

10 questions to ask any AI security vendor about data privacy

  1. Does any prompt, response or tool-call content leave my environment, and under what conditions?
  2. Do your detectors call a third-party LLM (OpenAI, Anthropic, Google or another) to make a verdict?
  3. Which of your services can run in my cloud account or data center, and on which Kubernetes distributions?
  4. What outbound connections does a self-hosted deployment need, and what data goes over each one?
  5. If I self-host, what data, if any, do you receive from my deployment?
  6. Where are logs stored at rest, who holds the keys, and how do I set retention and deletion?
  7. Which features are cloud-only today, and when will they be available self-hosted?
  8. Can you run with no outbound internet access? If so, how do updates arrive?
  9. How do you ship detector updates and security patches, and how do you disclose vulnerabilities in your own images?
  10. Can your red teaming reach applications that aren't exposed to the internet?

How Pillar runs inside your environment

Pillar Stack is Pillar's platform packaged as a Helm chart that runs inside your Kubernetes cluster. It's the same product our managed cloud runs, with your infrastructure underneath.

Two modes

  • Standalone mode: stateless runtime guardrails with no data persistence. Your applications or AI gateway send a prompt or response to the guardrail service in your cluster and get a verdict back. Nothing is stored. Use it for the most sensitive traffic, or where another system is already the audit record.
  • Full mode: the whole platform. Dashboard, sessions, findings, audit trail, AI discovery, red teaming and agentic endpoint posture, with data in your own PostgreSQL, Redis and Kafka, using TLS and your credentials.

What runs in your environment

CapabilitySelf-hostedHow it works
Runtime guardrails (prompt injection, jailbreak, PII, PCI, secrets, toxicity and more)YesDetection models run on GPU nodes in your cluster; prompt and file scanning APIs are served from your ingress. Detectors that use an LLM are pointed at a model endpoint in your own account, such as your Amazon Bedrock, or disabled
Dashboard, sessions, findings, audit trailYes (Full mode)Stored in your Postgres
SIEM and notification integrationsYesExport to your SIEM and ticketing tools
AI discovery and AI-BOMYesPillar's scanner runs as a job in your own GitHub Actions, GitLab CI or Azure DevOps pipeline and reports to the service in your cluster
Red teaming (Red Graph)YesRuns from your cluster against your applications, including internal ones that aren't exposed to the internet
Agentic endpoint postureYesYour MDM points the scanner at an endpoint gateway in your cluster. Scans collect security configuration data only
AI gateway integrationsYesLiteLLM, Amazon Bedrock and others connect to the guardrail service in your cluster
Multi-turn conversation analysis (/protect)YesAnalyzes whole conversations, not single messages, inside your cluster
Adaptive guardrailsNot yetCloud-only today, because it needs managed model serving

Feature status as of October 2026. Ask your Pillar contact for the current matrix.

Infrastructure

  • Kubernetes: Amazon EKS is the documented path, with Karpenter-managed node pools or EKS Auto Mode. Other Kubernetes environments are supported case by case, and Azure AKS, Google Cloud and bare-metal installation guides are on the roadmap.
  • Compute: at least three GPU-enabled nodes (the g6 family is recommended) for detection models, plus two general-purpose nodes.
  • Identity: IRSA or EKS Pod Identity for workloads; SSO and SCIM for users.
  • Delivery: Helm, or Flux for GitOps, with a Terraform module for IAM roles.
  • Production: run each service with at least three replicas and use your own clustered Postgres and Redis.
  • Network: Standalone mode stores nothing. Full mode needs outbound access to a short, documented list of domains for authentication and feature configuration; prompt content isn't sent to them. Sending telemetry to Pillar is opt-in. Whether a deployment counts as air-gapped depends on your definition. Pillar runs in production with no data flowing back to Pillar: images and detector updates are pulled into your environment, and detectors that use an LLM call an endpoint in your own account or are switched off. If your definition also rules out outbound authentication and feature-configuration services, we scope the deployment with you.
  • Updates: new detector models ship as images you pull and roll out on your own schedule.

A hybrid that keeps one policy. Many customers self-host the guardrails for regulated workloads and use Pillar's managed cloud for everything else. Policies, detectors and reporting work the same way in both, so the security team manages one program instead of two.

Pillar runs self-hosted in production at enterprise customers today, including a Fortune 300 manufacturer that runs it fully air-gapped. Pillar is SOC 2 Type II and ISO/IEC 27001:2022 certified (Trust Center).

Next steps

FAQs

What is on-premise AI security?

It's a deployment model in which the controls that discover AI assets, test them and inspect live AI traffic run on infrastructure the organization controls. Prompts, responses, tool calls, logs and findings stay inside its own boundary instead of going to the vendor's cloud.

Does an AI security vendor see my prompts?

A SaaS AI security vendor does. To detect prompt injection or data leakage, it must receive the full prompt, response and tool calls. A self-hosted deployment runs the same inspection inside your environment, so the vendor never receives that content.

Does an AI guardrail vendor become a GDPR sub-processor or a HIPAA business associate?

If it receives personal data or PHI from your AI traffic, generally yes. Under GDPR it processes personal data on your behalf. Under HIPAA, HHS treats a vendor that receives or stores ePHI as a business associate even if the data is encrypted. A self-hosted guardrail that never receives the data avoids that role. Confirm with counsel.

What is the difference between self-hosted, on-premise and air-gapped AI security?

Self-hosted means the vendor's software runs in infrastructure you control, usually your own cloud account. On-premise usually means your own data center. Air-gapped means the deployment has no outbound internet access, and updates are imported manually.

Which AI security platforms can be deployed on-premise?

Few platforms cover the whole AI lifecycle self-hosted. Pillar Security runs runtime guardrails, AI discovery, red teaming and agentic endpoint posture inside your own Kubernetes cluster. Before you choose any vendor, ask which features are cloud-only, what outbound access the deployment needs, and whether detectors call an external model.

What are the best self-hosted AI guardrails for banks?

Look for guardrails that run detection models on your own compute, store audit logs in your database, offer a stateless mode, and integrate with your AI gateway and SIEM. Pillar meets all four and adds AI discovery, red teaming and endpoint posture in the same deployment, which simplifies DORA, NYDFS and GLBA third-party risk reviews.

How can a bank use LLMs without sending customer data to a third party?

Host the model in your own cloud account or data center, or use a provider under your existing contract. Run the AI security layer (guardrails, logging, red teaming) inside your environment too. Otherwise the security tool becomes the third party that receives the data.

Which platforms provide runtime guardrails and full audit trails for healthcare AI applications?

Pillar provides runtime guardrails that mask PHI before it reaches a model and block unsafe responses, plus a full audit trail of every AI interaction. Self-hosted, both the inspection and the audit trail stay inside the health system's environment.

Is self-hosted AI security as good as SaaS?

The detection can be the same, but you take on infrastructure, GPU capacity and upgrades, and some features may reach the cloud first. Ask for a written feature matrix. Pillar's guardrails, multi-turn analysis, dashboard, discovery, red teaming and endpoint features all run self-hosted; adaptive guardrails are cloud-only today.

Can AI red teaming run against internal applications?

Yes, if the red-teaming engine runs inside your network. Self-hosted, Pillar's Red Graph runs from your cluster against applications that aren't exposed to the internet.

Subscribe and get the latest security updates

Back to blog

MAYBE YOU WILL FIND THIS INTERSTING AS WELL

Costrails: How AI Guardrails Cut LLM Costs, Not Just Risk

By

Dor Sarig

and

October 8, 2026

Blog
Best AI Security Frameworks for AI Agents in 2026: NIST AI RMF, OWASP, MITRE ATLAS, ISO 42001 and SAIL Compared

By

Dor Sarig

and

October 7, 2026

Guides
AI Gateway Guardrails in 2026: How to Secure LiteLLM, Kong, TrueFoundry and Agent Router (and 12 Guardrail Providers Compared)

By

Dor Sarig

and

October 4, 2026

Guides