Guide · AI governance
How to use LLMs securely in regulated industries
To use an LLM securely in a regulated industry, treat the provider like any outside service that handles regulated data. Sign the contract the regulation requires, check the provider's training, retention, human review and location terms, and send only the data the task needs. Four questions decide the controls: what data goes in, where it goes, what decision the output feeds, and what the model can reach.
Last reviewed
The question, as asked in r/cybersecurity: how are regulated organizations securely using LLMs other than Copilot? This guide starts from the rules organizations already follow when they hand regulated data to an outside service or use a tool in a regulated decision. For broker-dealers, FINRA confirmed that approach in June 2024: its rules are “intended to be technology neutral” and apply whether a firm builds a generative AI tool or uses a third party’s.
So the work is not to find a compliant model. It is to answer four questions about each use and let the answers pick the controls.
The four questions that decide the controls
What data goes into the prompt? This decides which contract you need before the first request.
- US health data: HHS says a cloud provider that creates, receives, maintains or transmits electronic protected health information for a covered entity is a business associate. That holds “even if the CSP processes or stores only encrypted ePHI and lacks an encryption key”. The two sides must sign a business associate agreement. The guidance is about cloud computing in general and never mentions language models. Reading it onto an LLM API that receives patient data is our inference, but it is hard to see which part of the definition an API would escape.
- Personal data under the GDPR: when the provider processes it on the organization’s behalf, GDPR Article 28 applies. The processing must be “governed by a contract” requiring the provider to act “only on documented instructions from the controller”. When the use is likely to be high risk to people, Article 35 requires a data protection impact assessment before processing starts. Its trigger names processing “using new technologies”.
- Data with no regulated content: none of these contracts are triggered, and the other three questions still apply.
Where does the data go, and how long does it stay? Four terms matter. Does the provider train on it, does it keep it, can a human read it, and where is it processed. The next section compares four providers.
What decision does the output feed? A summary a person checks is a different risk from a score that decides a loan. Under the EU AI Act, Article 6 and Annex III make a system high-risk if it is intended to evaluate a person’s creditworthiness (fraud detection is excluded), or to price their life or health insurance. Article 6(3) lifts a system out of that category when it poses no significant risk of harm and meets one of four conditions, such as doing only a narrow procedural or preparatory task. A system that profiles people never gets the exemption. A high-risk system’s deployer then has to “assign human oversight to natural persons who have the necessary competence, training and authority”. For Annex III systems those obligations apply from 2 December 2027. FINRA expects a firm that uses generative AI inside its supervision to cover “model risk management, data privacy and integrity, reliability and accuracy of the AI model” in its policies.
What can the model reach and do? In a chat window, the regulated data at stake is mostly what users paste in. A model connected to a document index or to tools can reach everything its connections can reach, which is where LLM data leakage and excessive agency come from. OWASP’s advice is to “enforce document- and chunk-level authorization inside the index query”. The model should see only documents the user asking could open themselves.
What the big providers say they do with your prompts
These are each provider’s own statements for its business offerings, read on 28 September 2026. They are vendor claims, they change, and the contract you sign is what binds.
| Trains on your data? | Keeps prompts? | Human review? | Where processed | |
|---|---|---|---|---|
| OpenAI API | Not by default | Up to 30 days to provide the service and identify abuse, or longer if the law requires; OpenAI lists some endpoints and features with different terms; zero data retention for eligible endpoints and qualifying uses | Access limited to authorized employees and specialized contractors, including to investigate abuse | Not compared here |
| Azure OpenAI (Microsoft) | Not without your permission or instruction; not available to OpenAI | Abuse monitoring can store flagged samples, unless you are approved for modified abuse monitoring; stateful features such as the Responses API store message history | Automated first, “with additional reviews by human reviewers as necessary” | Your chosen geography by default; “Global” deployments may process “in any geography” the model is deployed, and “DataZone” deployments anywhere in the data zone |
| Amazon Bedrock | Not compared here; model providers “don’t have access to” prompts and completions | You choose per Region: zero retention, the model’s default, or AWS review mode | Only for models that require it, and only once you opt into review mode: “Some model providers require Amazon to conduct human review as a condition of access” | Not compared here |
| Anthropic commercial | Not by default, except conversations a user sends as feedback, which are kept up to 5 years and may be used for training; Team and Enterprise owners can turn the feedback button off | Not compared here | Not compared here | Not compared here |
OpenAI also says it can sign a BAA for its API. Anthropic’s statement covers its commercial products and points consumer plans to a separate article. The no-training defaults described here are for business accounts, which is one more reason to route staff to those accounts rather than personal ones.
An order that works
- Find what is already in use. Staff using personal accounts are shadow AI, and none of the contracts or settings below reach them.
- Sort each use by the four questions. A drafting assistant that never sees customer data needs far less than a claims triage tool that reads medical records.
- Sign before you send. The BAA or GDPR processor contract comes before any regulated data reaches the provider. FINRA says a firm “should evaluate Gen AI tools prior to deploying them”.
- Set the data terms in the platform, not only on paper. Choose the retention mode, the deployment region and who can see logs. OWASP’s advice is to “technically enforce no-train/no-retain rather than policy text alone”.
- Send less. OWASP again: “send only task-required fields to external providers”, and “never store secrets, credentials, or regulated data in system prompts”.
- Keep your own record. Your prompts and outputs hold the regulated data now. OWASP advises you to “restrict and scrub logs and traces” before monitoring tools collect them, and, for regulated workloads, to feed “AI-aware audit logging into SIEM”.
- Match oversight to the decision. If the output feeds a high-risk decision, name the people who check it and give them the authority to overrule it.
What does not work
Encryption instead of a contract. HHS closed this off for health data: a provider holding only encrypted data it cannot decrypt is still a business associate, and the “conduit” exception covers “transmission-only services”. An LLM provider has to read the prompt to answer it, so it does more than transmit.
Treating “no training” as the whole answer. Training is one of four terms. A provider can promise not to train on prompts and still keep them for up to 30 days, as OpenAI’s API may, or show them to human reviewers, as Azure may when it flags abuse.
Assuming one set of data terms covers every model and deployment. On Bedrock, each model sets the minimum retention it accepts. Some can only be used if you let AWS keep prompts for human review; AWS names Claude Fable 5 and 5.1. Bedrock does not switch that on quietly: under zero retention or the default mode, those models are unavailable. Turning on review mode to reach one is a retention decision, and it belongs with whoever signed off the data terms. On Azure, the deployment type matters the same way: a Global deployment may process prompts in any geography where the model is deployed.
Self-hosting as an exit from the rules. Running an open weight model on your own servers takes the provider out of the data path, so the business associate agreement and processor questions change. The use still does not escape regulation. FINRA’s rules apply to tools a firm builds for itself as much as to third-party ones, and the EU AI Act’s high-risk duties follow the use, not the host.
What this guide does not settle
It does not make any provider compliant. OpenAI’s own wording is that it signs BAAs “in support of customers’ compliance”. Compliance stays with the organization, and it depends on how each tool is configured and used. The HHS and GDPR texts predate these tools. How they apply to LLMs is argued from their general definitions, and no regulator cited here has ruled on a specific LLM product.
Questions and answers
Can a hospital send patient data to a hosted LLM?
HHS says a cloud service provider that creates, receives or maintains electronic protected health information for a covered entity is a business associate and needs a business associate agreement, even if the data is encrypted. Only transmission-only "conduit" services are excluded. The guidance does not mention LLMs; applying it to an LLM API that receives patient data is our reading of its definition. OpenAI, for example, says it can sign a business associate agreement for its API Platform. Without that contract in place, patient data should not go into the prompt.
Does a promise not to train on your data make an LLM provider safe for regulated data?
No. Training is one of four terms to check. OpenAI says its API may keep inputs and outputs for up to 30 days to provide the service and identify abuse. Azure may send flagged samples to human reviewers. On Amazon Bedrock, some models can only be used if you let AWS keep prompts for human review.
Sources
- Guidance on HIPAA & Cloud ComputingUS Department of Health and Human Services, Office for Civil Rights
- Regulation (EU) 2016/679 (GDPR), Article 28European Union, 27 Apr 2016
- Regulatory Notice 24-09, FINRA Reminds Members of Regulatory Obligations When Using Generative Artificial Intelligence and Large Language ModelsFINRA, 27 Jun 2024
- Regulation (EU) 2024/1689 (EU AI Act), Articles 6, 26 and 113 and Annex III (consolidated 27.07.2026)European Union, 12 Jul 2024
- LLM02:2026 Sensitive Information DisclosureOWASP GenAI Security Project, 1 Jan 2026
- Enterprise privacy at OpenAIOpenAI
- Data, privacy, and security for Foundry Models sold by Azure in Microsoft FoundryMicrosoft
- Data protection, Amazon Bedrock User GuideAmazon Web Services
- Data retention, Amazon Bedrock User GuideAmazon Web Services
- Is my data used for model training?Anthropic, 18 Aug 2026