Guide · AI security
How to detect shadow AI
To detect shadow AI, check several places it leaves traces. Proxy and firewall logs show visits to AI services. Identity-provider consents show AI apps granted access to company data. Secret scanning finds AI API keys in code, cloud audit logs record calls to AI services, and network scans find self-hosted models. Each source misses something. Network and endpoint tools see nothing from a personal device on a home network.
Last reviewed
Finding shadow AI is mostly a matter of reading records you can already get: network logs, the identity provider, code repositories and cloud accounts. Some need switching on first. Each one shows a different part of the picture, and each has a blind spot.
Where shadow AI leaves traces
| Where to look | What it shows | What it misses |
|---|---|---|
| Proxy and firewall logs | Which AI services are reached from your network, and from which users or IP addresses | Prompt content, personal devices off the network, AI inside apps you already approved |
| Endpoint and browser tools on managed devices | Visits to AI sites, and sensitive data pasted or uploaded to them | Unmanaged devices, and devices or browsers the tool does not support |
| Identity provider app consents | AI apps that users signed in to with work accounts and gave access to mail, files or calendars | Tools used without a work sign-in |
| Code repositories | API keys for AI providers committed to code | Keys kept out of the repository, and repositories nobody scans |
| Cloud audit logs | Calls to managed AI services in your cloud accounts, and who made them | Accounts you do not know about, and events not logged by default |
| Network scans | Self-hosted model servers listening on the network | Servers on non-standard ports or bound to one machine |
Work through the sources in this order
-
Start with traffic you already log. IBM’s advice is to “implement network monitoring tools to track application usage”. In practice this means matching proxy or firewall logs against a catalog of known AI services. Microsoft’s Defender for Cloud Apps is one example. It matches uploaded firewall and proxy logs against a catalog of over 31,000 cloud apps, with categories for “Generative AI” apps, model provider APIs and MCP servers. Its shadow AI guide says to filter for the Generative AI category first, and to do this “before taking any blocking actions”. The result is a list of AI services in use, and a ranking of the top users and source IP addresses. Not every log names the user: Microsoft’s table of supported log formats shows several, including Check Point and Cisco ASA, with no user name.
-
Check which AI apps hold access to company data. An AI meeting assistant or writing tool that a user signs in to with a work account shows up in the identity provider. Google Workspace lists “Accessed apps”, with the number of users and the OAuth scopes each app uses, such as Gmail or Drive. Microsoft Entra shows the permissions each app holds, both those granted for the whole organization and those granted to a specific user or group. This is the one source that says what an AI app can read, not just that someone visited it.
-
Scan code for AI API keys. A team calling a model provider from its own code needs an API key, and a key committed to a repository shows that the team is using the provider. GitHub’s secret scanning recognizes API keys from Anthropic, OpenAI and other providers. It runs free on public repositories. Private and internal repositories need GitHub’s paid Secret Protection.
-
Read your cloud audit logs. Teams that build on managed AI services leave records in the cloud account. AWS CloudTrail records calls to Amazon Bedrock with who made each one, from which IP address, and when. Some of these logs are off by default: CloudTrail does not log high-volume “data events” unless you turn them on, and Google Cloud’s Vertex AI Data Access audit logs are “disabled by default”. Turn them on before relying on them.
-
Look for self-hosted models. A model running locally with a tool such as Ollama answers prompts on the machine itself. Traffic logs may see the model being downloaded, but not the prompts, unless the user switches to Ollama’s cloud-hosted models. Ollama listens on port 11434 and, by default, only on the machine it runs on. When someone changes that setting, the server becomes reachable over the network. In 2025 Cisco researchers used the Shodan search engine to find 1,139 Ollama servers exposed to the internet, 214 of them serving live models. They warn that the method looks for default ports such as 11434, so servers on other ports “are likely to evade detection entirely”. A scan of your own address ranges for port 11434 finds servers exposed on that port, not ones moved to another. The risks of running LLMs locally cover what else to check on those machines.
-
Write down what you find. Discovery that is not recorded has to be repeated. OWASP’s governance checklist says to “Catalog existing AI services, tools, and owners”, include AI components in the software bill of materials, and “Create an AI solution onboarding process”. NIST’s AI Risk Management Framework asks for “Mechanisms … to inventory AI systems”. An AI bill of materials is one format for that inventory.
What detection will not find
Personal devices and personal accounts. Traffic logs see only traffic that passes through them. Microsoft’s endpoint integration exists to “extend cloud discovery capabilities beyond your corporate network”. A personal phone on a home network sends nothing to either. The UK NCSC adds a subtler gap. Without breaking encrypted connections, network monitoring “will not be able to identify unapproved use of approved cloud services (i.e. personal accounts)”. A personal ChatGPT account and a company one reach the same domain.
AI that arrives inside approved software. OWASP’s description of shadow AI includes “third-party applications that introduce LLM features via updates or upgrades”. The app is already approved, so a check against a list of approved apps passes it, even after it gains an AI feature.
Anything the catalog does not know. Microsoft says its discovery “cannot discover apps that aren’t in the catalog” by default. A new AI tool is invisible to catalog matching until someone adds it.
What was sent. Knowing that someone used an AI service is not knowing what they pasted into it. Traffic logs show the service, not the content. Microsoft’s Global Secure Access logs prompt content through a separate feature, Generative AI Insights, which decrypts and inspects the traffic. Microsoft Purview does not see sensitive data shared with third-party AI sites from a device that is not onboarded to it.
Ask people, and make it safe to answer
People can also be asked. The NCSC warns that staff “will be reluctant to come forward if they fear they (or other members of staff) will be reprimanded”, and that “a poor security culture means you’re much less likely to detect shadow IT”. So a survey, an amnesty on unapproved tools or a quick route to approval can surface use that no log records, as long as nobody is punished for answering. The answers also say why people went around the approved tools.
Questions and answers
Can a firewall or proxy see what employees type into ChatGPT?
Not from ordinary traffic logs. Those show that someone reached an AI service, not what they sent. Microsoft's Global Secure Access can log prompt content, but only by decrypting and inspecting the traffic. Microsoft Purview can see sensitive data shared with AI sites, but only from devices onboarded to it. Neither sees a personal phone on a home network.
Should we block AI tools once we find them?
Find out what people use them for first. Microsoft's own deployment guide puts discovery and monitoring "before taking any blocking actions". If blocking comes with blame, the UK NCSC's warning applies: staff who fear being reprimanded "will be reluctant to come forward", which makes shadow IT harder to detect.
Sources
- LLM Applications Cybersecurity and Governance Checklist v1.1OWASP Gen AI Security Project, 31 Oct 2024
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1NIST, Jan 2023
- What is shadow AI?IBM, 25 Oct 2024
- Shadow ITUK National Cyber Security Centre, 27 Jul 2023
- Cloud app discovery overviewMicrosoft
- View discovered apps on the Cloud discovery dashboardMicrosoft
- Cloud app catalog and risk scoresMicrosoft
- Shadow AI discovery in Global Secure AccessMicrosoft
- Considerations for Data Security Posture Management for AI (classic)Microsoft
- Step 1: Discover AI appsMicrosoft
- Control which third-party & internal apps access Google Workspace dataGoogle
- Review permissions granted to enterprise applicationsMicrosoft
- Supported secret scanning patternsGitHub
- Monitor Amazon Bedrock API calls using CloudTrailAmazon Web Services
- Vertex AI audit loggingGoogle Cloud
- Ollama FAQOllama
- Detecting Exposed LLM Servers: A Shodan Case Study on OllamaCisco, 1 Sep 2025