API Key Sprawl in SaaS Platforms and Its Role in Breaches
Untracked API keys become permanent breach entry points that traditional security tools can't find.
Summary
Untracked API keys become permanent breach entry points that traditional security tools can't find.
API key sprawl in SaaS environments is not a tidiness problem someone forgot to clean up. It is a structural condition built into how modern software talks to itself, where untracked, over-permissioned, long-lived keys sit around as permanent attack surfaces that ordinary security tools were never designed to see, let alone close.
That's the deal software vendors made a long time ago: keys are cheap to issue, fast to check, and they don't ask questions. Whoever holds the token gets in, full stop. No MFA prompt, no re-authentication, no check on whether the behavior looks like the account it claims to be. Once a key is handed out, it works independently of whatever SSO or MFA setup a company has spent years and budget building. The identity infrastructure that gates human logins simply doesn't apply once you're a machine holding a string of characters.
Broad permissions granted to integrations compound this. When an admin authorizes an integration and the app asks for broad scopes, the app usually gets them, inheriting the same level of access as the person who clicked "allow". Multiply that across a workplace: enterprises now manage roughly 490 cloud apps, many of them unsanctioned, and each integration creates a non-human identity credential that traditional identity management doesn't track. None of this would matter much if the traffic were small. The traffic is not small. A credential built for machines is being governed, if you can call it that, by processes built for people. API keys function as the connective tissue of SaaS environments, with every integration, automation, and third-party service depending on them. SaaS-to-SaaS data moves 10× faster than human-to-SaaS interactions, according to Obsidian's network data, and the velocity that makes keys valuable also amplifies breach impact.
Where API keys end up
Sprawl behaves a lot like shadow IT always has. Someone pastes a key into an environment variable to hit a deadline, someone else drops a token into a Slack thread to unblock a teammate. Each shortcut makes sense in the moment. Stacked across a company over years, they add up to a credential inventory that nobody actually owns.
Credentials live across five surfaces simultaneously. Deployment platforms like Vercel, Netlify, and Render collect keys that get pasted in once and then forgotten, each one managed on its own island. CI/CD pipelines, GitHub Actions, GitLab CI, Jenkins, carry outsized risk because a pipeline needs broad deployment access to function; in the Shai-Hulud 2 supply chain attack, 59% of the compromised machines turned out to be CI/CD runners GitGuardian State of Secrets Sprawl 2026. Infrastructure tooling such as Terraform lets credentials drift between workspaces with rotation handled unevenly, if at all. And then there's the everyday sprawl: developer laptops,.env files, shell history, and the Slack messages, Jira tickets, and Notion pages where a key got shared "just this once" and stayed there.
The public-GitHub obsession in most security tooling misses where the real exposure sits. Worse, 28% of secrets incidents start outside code entirely, in Slack, Jira, Confluence, and those non-code leaks turn out to be 13% more likely to get flagged as critical than the ones found in a repo GitGuardian State of Secrets Sprawl 2026. That invisibility turns leaked keys into breaches at the scale the next section lays out. Internal repos are 6× more likely than public ones to contain secrets, so the "public GitHub" focus in tooling understates the internal exposure, per the GitGuardian 2026 report GitGuardian State of Secrets Sprawl 2026. Without a complete inventory, rotation is guesswork, blast radius during an incident is unknowable, and revocation of stale access is impossible, according to Doppler.
The scale of what has already leaked
The numbers here aren't estimates dressed up to sound alarming. GitGuardian combed billions of commits across public GitHub and counted 28.65 million new hardcoded secrets exposed in 2025 alone, a 34% jump from the year before and the biggest single-year increase the firm has ever recorded. Since 2021, leaked secrets have grown 152%, while the population of developers on GitHub grew only 98% over that same stretch. Secrets are multiplying faster than the people who write them, which is a strange thing to say out loud and an even stranger thing to actually be true.
SpyCloud got at the same problem from a completely different angle, mining breach and malware data rather than code repositories, and came back with 18.1 million exposed API keys and tokens recaptured in 2025.
Then there's the problem of datasets containing exposed secrets. In February 2025, Truffle Security went looking through Common Crawl's public dataset, roughly 400 terabytes of archived web pages covering more than 2.67 billion pages, and found close to 12,000 valid API keys and credentials sitting inside it, including working keys for AWS and MailChimp 12,000 API Keys Exposed in LLM Training Data Leak!. Common Crawl is one of the datasets widely used to train large language models, so those keys weren't just exposed to human attackers 12,000 API Keys Exposed in LLM Training Data Leak!. They were baked into training data.
Across 6,943 systems examined, GitGuardian found 294,842 secret occurrences tied to 33,185 unique secrets, with each live secret showing up in about eight different places on the same machine. One leak rarely stays one leak. It replicates itself across environment files, config files, and scripts until a single compromised credential looks less like a needle and more like a haystack made entirely of needles.
Why so many leaked keys remain exploitable years later
Detection is not the bottleneck here. Remediation is. GitGuardian's research found that 64% of valid secrets leaked back in 2022 were still active and exploitable four years later. That's not a blind spot, that's a policy failure sustained over half a decade.
Most organizations only revoke keys once a breach has already been found; proactive rotation before something goes wrong remains rare. Stale integrations make this worse, not better. Permissions granted for a specific project tend to outlive the project itself, so broad access sits attached to an app nobody's thought about in a year, watched by nobody. Broad access plus zero oversight is about as toxic a combination as exists in a security stack.
The longer a key sits valid, the harder it gets to trace an active exploitation back to the moment it was first exposed. Investigations get slower exactly when speed matters most. Longevity is the one property that takes a single leaked credential and turns it into a standing, indefinite entry point. And right now, AI tooling is filling that entry point with new candidates faster than any team can reasonably track. API keys are designed to be long-lived or indefinitely applied with no built-in session duration, so an attacker exploiting a key months after initial exposure faces no automatic expiry, according to NordicAPIs.
How AI tooling is accelerating the rate of credential exposure
GitGuardian counted 1,275,105 leaked secrets tied to AI services in 2025, up 81% from the year before. Eight of the ten fastest-growing categories of leaked secrets, year over year, now trace back to AI tooling specifically, which makes AI adoption the leading driver of new credential exposure rather than a side effect of it.
Some of that comes down to how the code gets written. AI-assisted commits leak secrets at roughly twice the rate of commits written by hand. It isn't the flashy model providers doing the damage, either. MCP configuration files turned out to be a specific offender: GitGuardian found 24,008 unique secrets sitting in MCP-related configs on public GitHub in 2025, with 2,117 of them confirmed still valid, a problem made worse by official documentation that, for a while, actually walked developers toward hardcoding credentials directly into those files.
AI agents themselves are a new kind of surface altogether. They run continuously, request secrets on the fly, and act without a human clicking anything, which makes them both a fresh source of sprawl and a category no inventory tool was ever built to watch. The same wave of AI adoption that's expanding how many SaaS tools talk to each other is, at the same time, generating credentials faster than any manual cleanup process could ever keep pace with. LLM infrastructure, including RAG pipelines, orchestration layers, and vector storage, leaked credentials 5× faster than core model providers, showing the surrounding ecosystem is the bigger risk, according to GitGuardian via daily.dev.
How leaked keys translate into actual breaches, four documented cases
64% of valid secrets leaked in 2022 were still active and exploitable four years later, showing this is a remediation gap.
The US Treasury breach in late 2024 didn't require anything clever. Attackers used a leaked API key for BeyondTrust's Remote Support SaaS platform, and that single exposed credential walked straight past a security budget most organizations would envy, landing them inside Treasury systems.
The Common Crawl exposure in March 2025 showed the same failure mode running through a newer pipe 12,000 API Keys Exposed in LLM Training Data Leak!. Close to 12,000 valid API keys sat inside a dataset used to train large language models, meaning credentials traveled straight into the AI supply chain through a route no secrets scanner was watching for 12,000 API Keys Exposed in LLM Training Data Leak!.
The Salesloft/Drift breach in August 2025 is the one with the most moving parts. Attackers harvested OAuth tokens, mostly for Salesforce integrations but also touching Google Workspace and Slack, from the Drift chatbot platform, then used those tokens to pull contact records, account records, and support case data out of hundreds of customer organizations. Both the initial compromise and the exploitation window occurred within August 2025, roughly August 8-18, showing how longevity of valid tokens extends an attacker's operational window. One chatbot vendor's OAuth grant, multiplied by every company plugged into it.
Vercel's incident in April 2026 followed the same script with a different cast. A third-party AI tool called Context.ai got compromised through an abused Google Workspace OAuth grant, and attackers walked away with API keys, npm tokens, database credentials, and GitHub tokens, later listing the stolen data for sale on BreachForums at $2 million.
Lining these four up shows the pattern repeats without variation: no exotic exploit, just a static credential, usually inherited through some integration nobody was actively watching. Obsidian researchers found this attack's blast radius was 10× greater than previous incidents where attackers infiltrated Salesforce directly, with the integration chain multiplying exposure to more than 700 companies, according to Obsidian Security and CyberTech Intelligence.
Why broken authentication and behavior-based abuse dominate the breach data
An analysis of 60 publicly disclosed API breaches in 2025 found broken authentication behind 52% of them, the single largest cause by a wide margin SaaS API Security Incidents: What 2026 Breach Data Shows - Security B…. Add unsafe consumption of third-party APIs, breaches that start with a connected vendor or integration rather than the company's own code, at 27%, and those two categories alone account for nearly 80% of API breaches, both of them direct downstream effects of untracked, over-permissioned credentials sitting inside integration chains SaaS API Security Incidents: What 2026 Breach Data Shows - Security B….
The type of attack is shifting too. Behavior-based attacks, meaning abuse of legitimate workflows rather than classic exploitation of broken code, made up 61% of API attacks in 2025, up sharply from 30% just a year earlier SaaS API Security Incidents: What 2026 Breach Data Shows - Security B…. That shift changes what defenses need to catch, since scanners built to flag malformed requests or known signatures have nothing broken to detect. Behavior-based abuse doesn't trip a scanner looking for malformed requests or known signatures, because there's no broken code to find. The attacker is just using a real key the way it was designed to be used, only for the wrong purpose.
Meanwhile the volume keeps climbing. Average daily API attacks per organization rose 113% year over year, according to Akamai's State of the Internet report, so scale and sophistication are both moving in the wrong direction at once. And detection lags badly behind all of it: 38% of organizations only found out about their API breaches after someone outside the company reported it, not through internal detection, and 47% of API endpoints sat exposed for six months or longer before anyone noticed. Compromised credentials served as an initial access vector in 22% of breaches, and stolen credentials appeared in 88% of Basic Web Application attacks, according to Verizon's DBIR.
Why conventional security controls were not built to close this gap
SSO and MFA do one job well: they protect the moment a human logs in. OAuth tokens and API keys skip that moment entirely, and once one is issued it keeps working until somebody actively revokes it. Cloud access security brokers were built to watch traffic between a user and an app, but they have nothing to say about app-to-app traffic, which is exactly where API keys spend their time.
Static inventories don't fare much better. They capture a snapshot of what exists at one moment, but they can't tell when a perfectly legitimate credential is being misused, because there's no broken code involved, just a workflow running outside the bounds it was meant to stay inside. Even companies that have invested heavily in security tooling admit the gap: 56% of enterprises say they lack full visibility into their own API data flows, and only 10% of organizations have any kind of posture governance strategy specifically for API security, meaning nearly everyone has detection tools sitting next to an empty space where a governance framework should be.
Every access model built for humans assumes a session, a login event, and a person who can be challenged mid-action. Agents and API keys generate none of those signals. Plenty of organizations log API activity diligently and still have no runtime mechanism to stop a violation before it finishes happening. They're good at writing history and bad at preventing it. What's missing across the board isn't more scanning of code repositories. It's real-time behavioral monitoring paired with policy controls that can actually enforce something, not just record it.
What effective API key governance requires in practice
A complete inventory has to come first. Without knowing every key, every integration, and every surface where a credential might live, rotation and revocation are both just guesswork dressed up as process. Keys should also be scoped to the minimum an integration actually needs at the moment they're issued, not inherited wholesale from whichever admin clicked "allow," since inherited admin access is precisely the structural flaw that made the Salesloft/Drift breach spread as far as it did.
Rotation needs to stop being a post-breach reaction and become automatic, tied to short time-to-live windows that shut the exploitability gap responsible for 64% of 2022's leaked secrets still working in 2026. That means scanning continuously across all five surfaces, not just code repositories: CI/CD pipelines, deployment platforms, collaboration tools, infrastructure tooling, and AI agent configurations all need coverage, because leaving one uncovered just relocates the sprawl rather than solving it.
Behavioral monitoring belongs in the stack as its own layer. With 61% of API attacks in 2025 built around behavior rather than exploitation, detection has to watch for anomalies in how legitimate credentials get used, unusual data exports, logins from unexpected IP ranges, access patterns that drift from what's normal, rather than waiting for a malformed request to show up. AI agents need their own governance track entirely, since they inherit human-scale permissions but act at machine speed with no login event for anyone to notice.
Owning a secrets manager isn't the deciding factor. Most do. For most organizations right now, the gap between having tools and having governance, rather than one control plane giving real-time visibility, behavioral detection, and enforceable rules across every surface at once, is exactly where the breaches keep coming from.
Sources
- Why secrets sprawl is your biggest risk in 2026
- SaaS API Security Incidents: What 2026 Breach Data Shows - Security Boulevard
- What is API Security? Protecting the Hidden Layer Between SaaS Apps
- 12,000 API Keys Exposed in LLM Training Data Leak!
- The State of Secrets Sprawl 2026 | GitGuardian Annual Report
- 64% of Leaked Secrets Still Work Years Later
- API Key Management: Risks & Best Practices | Akeyless