Where do secrets leak?
Secrets leak wherever code or build output travels, and most leaks happen outside the main branch that people review. The places a scanner has to cover:
- Source code and config files. A key pasted into
config.py, a.envfile committed by mistake, a test fixture with a real token. These are the easiest to catch in secure code review, and the easiest to miss in a large diff. - Git history. Deleting the line in a later commit does not remove it. Every clone and fork still holds the old commit, and
git log -pshows it. - CI logs. A debug step that prints the environment, a tool run with verbose flags, a failed command that echoes its arguments. CI platforms mask values they know are secrets, but not values they were never told about; GitHub Actions provides
::add-mask::for those. - Container images. The Docker documentation warns that build arguments and environment variables "persist in the final image." A file copied in one layer and deleted in the next is still in the earlier layer, and
docker historyshowsARGandENVvalues. - JavaScript bundles. Build tools inline environment variables into client code when told to, such as variables prefixed
NEXT_PUBLIC_orVITE_. Anything placed there ships to every browser, and source maps can reveal more. - Mobile apps. Strings compiled into an Android APK or iOS app can be extracted with standard decompilers. Obfuscation slows this down; it does not hide a key.
- Everywhere else. Tickets, chat messages, wiki pages, Postman collections and public gists.
A leaked key widens your attack surface to whatever that key can reach, which for a cloud or CI token can be most of your infrastructure. CWE-798, Use of Hard-coded Credentials, is the weakness class for the code side of this.
How does secrets detection work?
Scanners combine three methods, each covering the others' gaps.
1. Provider patterns. Many providers issue keys with a fixed prefix and length, so a regular expression can match them with few false positives. AWS access key IDs, for instance, start with a four-letter prefix such as AKIA followed by 16 uppercase letters and digits. GitHub's own tokens use prefixes like ghp_ for this reason. Pattern rules are precise but only find formats someone has written a rule for.
2. Entropy. Random-looking strings are suspicious. A scanner measures Shannon entropy, a score of how unpredictable the characters are, and flags high-entropy strings assigned to names like secret, token or password. Gitleaks, for example, lets a rule require a minimum entropy for the matched group. A minimal version of the check:
import math
from collections import Counter
def shannon_entropy(s: str) -> float:
counts = Counter(s)
return -sum((n / len(s)) * math.log2(n / len(s)) for n in counts.values())
candidates = {
"api_token": "EXAMPLEONLYx9Qm2Lr7Vt4Kp8Zw3Nc6Hb1Jd5", # placeholder, not a real key
"log_level": "information",
}
for name, value in candidates.items():
flagged = len(value) >= 20 and shannon_entropy(value) > 4.0
print(f"{name}: entropy={shannon_entropy(value):.2f} flagged={flagged}")api_token: entropy=4.97 flagged=True
log_level: entropy=2.91 flagged=FalseEntropy catches secrets with no known format, such as internal service passwords, and also flags hashes, UUIDs and test fixtures. It is the noisiest method.
3. Verification. The scanner takes a candidate and asks the provider whether it works, using a harmless read-only call. TruffleHog, for example, checks AWS credentials with GetCallerIdentity and labels results as verified, unverified or unknown. A verified finding is a live credential and goes to the top of the queue. Verification sends the candidate to the provider, so only run it where your policy allows, and never "verify" a secret you found in someone else's system.
Pre-commit vs CI vs history scanning
The three scan points catch leaks at different costs, and a mature setup uses all three.
| Where | When it runs | Strength | Weakness |
|---|---|---|---|
| Pre-commit hook | On the developer machine before each commit | Secret never enters history | Optional; skipped with --no-verify or never installed |
| Push or CI check | On push or pull request | Enforced for everyone | Secret is already in a local commit that must be rewritten |
| History scan | On a schedule or once per repository | Finds leaks already sitting in old commits | Findings may already be exploited |
A pre-commit hook with an open-source scanner looks like this:
# .pre-commit-config.yaml
repos:
- repo: https://github.com/gitleaks/gitleaks
rev: v8.30.1
hooks:
- id: gitleaksIn CI, scan only the commits in the pull request so the job stays fast, and run a full-history scan separately, including archived repositories. The shift-left security page shows a PR workflow that includes this step.
How does GitHub push protection work?
Push protection scans pushes for supported secret types and rejects the push before the secret reaches GitHub. As GitHub documents it in September 2026, it exists at two levels:
- Push protection for users is enabled by default on every GitHub.com account and blocks you from pushing supported secrets to public repositories.
- Repository push protection is disabled by default and must be turned on by an administrator. It requires GitHub Secret Protection and covers the command line, the web UI, file uploads and the REST API.
When a push is blocked, the developer sees a message explaining which secret was found and where, removes it from the commit, and pushes again. A user with write access can bypass the block by choosing a reason (used in tests, false positive, or fix later), and organizations can require that bypasses go through designated reviewers. GitHub notes that push protection only supports the most recent token versions it can identify with confidence, to avoid blocking on likely false positives, and passwords found by its AI detection are not covered. Secret scanning of public repositories also runs automatically across full history, and GitHub notifies partner providers of their leaked tokens so they can act on them. Push protection is a backstop for supported formats, not a replacement for scanning your own history and artifacts.
What do you do when a secret leaks?
Revoke first, clean up second. Rewriting history does nothing about a copy someone already pulled, and GitHub's guidance on removing sensitive data says the first step for a leaked credential is to revoke or rotate it. A response procedure:
- Revoke or rotate the credential at the provider immediately, and deploy the replacement. Assume it was copied the moment it was pushed to a public repository.
- Check the provider's logs for use of the old credential between the leak and the revocation: API calls, new users or keys created, unusual source IPs.
- Find every copy. Search other repositories, CI logs, container registries and chat for the same value.
- Remove it from history once it is dead, so it stops generating alerts and confusion. GitHub documents
git-filter-repoversion 2.47 or later for this:
# Replace every listed value in all history, then force-push
git-filter-repo --sensitive-data-removal --replace-text ../leaked-values.txt
git push --force --mirror origin- Handle what the rewrite cannot reach. Forks keep the old commits, collaborators' clones can push them back, and cached views and pull request references on GitHub need a request to GitHub Support.
- Fix the cause. Move the value into a secrets manager or the CI platform's secret store, add a pre-commit hook, and add the pattern to your scanner if it was missed.
The OWASP Secrets Management Cheat Sheet describes the same order: immediate revocation, then rotation, then removal from code and logs. Treat a leaked CI token as a pipeline compromise, and review the rest of your CI/CD pipeline security with that in mind.
[ Sources ]
Written by Parameter · Last reviewed

