Parameter

Secrets detection

Also known as

  • Secret scanning
  • Credential scanning

Secrets detection is the automated search of source code, git history, CI logs, container images and build artifacts for credentials such as API keys, tokens, private keys and database passwords, using provider patterns, entropy checks and live verification, so exposed secrets are found and revoked before an attacker uses them.

Last reviewed

Where do secrets leak?

Secrets leak wherever code or build output travels, and most leaks happen outside the main branch that people review. The places a scanner has to cover:

  • Source code and config files. A key pasted into config.py, a .env file committed by mistake, a test fixture with a real token. These are the easiest to catch in secure code review, and the easiest to miss in a large diff.
  • Git history. Deleting the line in a later commit does not remove it. Every clone and fork still holds the old commit, and git log -p shows it.
  • CI logs. A debug step that prints the environment, a tool run with verbose flags, a failed command that echoes its arguments. CI platforms mask values they know are secrets, but not values they were never told about; GitHub Actions provides ::add-mask:: for those.
  • Container images. The Docker documentation warns that build arguments and environment variables "persist in the final image." A file copied in one layer and deleted in the next is still in the earlier layer, and docker history shows ARG and ENV values.
  • JavaScript bundles. Build tools inline environment variables into client code when told to, such as variables prefixed NEXT_PUBLIC_ or VITE_. Anything placed there ships to every browser, and source maps can reveal more.
  • Mobile apps. Strings compiled into an Android APK or iOS app can be extracted with standard decompilers. Obfuscation slows this down; it does not hide a key.
  • Everywhere else. Tickets, chat messages, wiki pages, Postman collections and public gists.

A leaked key widens your attack surface to whatever that key can reach, which for a cloud or CI token can be most of your infrastructure. CWE-798, Use of Hard-coded Credentials, is the weakness class for the code side of this.

How does secrets detection work?

Scanners combine three methods, each covering the others' gaps.

1. Provider patterns. Many providers issue keys with a fixed prefix and length, so a regular expression can match them with few false positives. AWS access key IDs, for instance, start with a four-letter prefix such as AKIA followed by 16 uppercase letters and digits. GitHub's own tokens use prefixes like ghp_ for this reason. Pattern rules are precise but only find formats someone has written a rule for.

2. Entropy. Random-looking strings are suspicious. A scanner measures Shannon entropy, a score of how unpredictable the characters are, and flags high-entropy strings assigned to names like secret, token or password. Gitleaks, for example, lets a rule require a minimum entropy for the matched group. A minimal version of the check:

import math
from collections import Counter

def shannon_entropy(s: str) -> float:
    counts = Counter(s)
    return -sum((n / len(s)) * math.log2(n / len(s)) for n in counts.values())

candidates = {
    "api_token": "EXAMPLEONLYx9Qm2Lr7Vt4Kp8Zw3Nc6Hb1Jd5",  # placeholder, not a real key
    "log_level": "information",
}
for name, value in candidates.items():
    flagged = len(value) >= 20 and shannon_entropy(value) > 4.0
    print(f"{name}: entropy={shannon_entropy(value):.2f} flagged={flagged}")
api_token: entropy=4.97 flagged=True
log_level: entropy=2.91 flagged=False

Entropy catches secrets with no known format, such as internal service passwords, and also flags hashes, UUIDs and test fixtures. It is the noisiest method.

3. Verification. The scanner takes a candidate and asks the provider whether it works, using a harmless read-only call. TruffleHog, for example, checks AWS credentials with GetCallerIdentity and labels results as verified, unverified or unknown. A verified finding is a live credential and goes to the top of the queue. Verification sends the candidate to the provider, so only run it where your policy allows, and never "verify" a secret you found in someone else's system.

Pre-commit vs CI vs history scanning

The three scan points catch leaks at different costs, and a mature setup uses all three.

WhereWhen it runsStrengthWeakness
Pre-commit hookOn the developer machine before each commitSecret never enters historyOptional; skipped with --no-verify or never installed
Push or CI checkOn push or pull requestEnforced for everyoneSecret is already in a local commit that must be rewritten
History scanOn a schedule or once per repositoryFinds leaks already sitting in old commitsFindings may already be exploited

A pre-commit hook with an open-source scanner looks like this:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.30.1
    hooks:
      - id: gitleaks

In CI, scan only the commits in the pull request so the job stays fast, and run a full-history scan separately, including archived repositories. The shift-left security page shows a PR workflow that includes this step.

How does GitHub push protection work?

Push protection scans pushes for supported secret types and rejects the push before the secret reaches GitHub. As GitHub documents it in September 2026, it exists at two levels:

  • Push protection for users is enabled by default on every GitHub.com account and blocks you from pushing supported secrets to public repositories.
  • Repository push protection is disabled by default and must be turned on by an administrator. It requires GitHub Secret Protection and covers the command line, the web UI, file uploads and the REST API.

When a push is blocked, the developer sees a message explaining which secret was found and where, removes it from the commit, and pushes again. A user with write access can bypass the block by choosing a reason (used in tests, false positive, or fix later), and organizations can require that bypasses go through designated reviewers. GitHub notes that push protection only supports the most recent token versions it can identify with confidence, to avoid blocking on likely false positives, and passwords found by its AI detection are not covered. Secret scanning of public repositories also runs automatically across full history, and GitHub notifies partner providers of their leaked tokens so they can act on them. Push protection is a backstop for supported formats, not a replacement for scanning your own history and artifacts.

What do you do when a secret leaks?

Revoke first, clean up second. Rewriting history does nothing about a copy someone already pulled, and GitHub's guidance on removing sensitive data says the first step for a leaked credential is to revoke or rotate it. A response procedure:

  1. Revoke or rotate the credential at the provider immediately, and deploy the replacement. Assume it was copied the moment it was pushed to a public repository.
  2. Check the provider's logs for use of the old credential between the leak and the revocation: API calls, new users or keys created, unusual source IPs.
  3. Find every copy. Search other repositories, CI logs, container registries and chat for the same value.
  4. Remove it from history once it is dead, so it stops generating alerts and confusion. GitHub documents git-filter-repo version 2.47 or later for this:
# Replace every listed value in all history, then force-push
git-filter-repo --sensitive-data-removal --replace-text ../leaked-values.txt
git push --force --mirror origin
  1. Handle what the rewrite cannot reach. Forks keep the old commits, collaborators' clones can push them back, and cached views and pull request references on GitHub need a request to GitHub Support.
  2. Fix the cause. Move the value into a secrets manager or the CI platform's secret store, add a pre-commit hook, and add the pattern to your scanner if it was missed.

The OWASP Secrets Management Cheat Sheet describes the same order: immediate revocation, then rotation, then removal from code and logs. Treat a leaked CI token as a pipeline compromise, and review the rest of your CI/CD pipeline security with that in mind.

Written by Parameter · Last reviewed

[ related terms ]

Related terms.

Shift-left security

Shift-left security is the practice of moving security checks earlier in software development, into design, the developer's editor and the pull request, so threat models, static analysis, dependency checks and secrets scanning catch flaws before code merges, while a fix is still a small edit to the author's own change.

Static application security testing (SAST)

Static application security testing (SAST) is automated analysis of source code, bytecode or binaries, without running the application, that traces untrusted input to dangerous operations such as SQL queries, shell commands and HTML output, and reports the file and line where an injection or similar flaw could occur.

CI/CD pipeline security

CI/CD pipeline security is the practice of controlling who and what can change, trigger and run your build and deployment pipelines, and what those pipelines can reach, so an attacker who gets into a branch, a third-party action or a runner cannot use the pipeline's secrets and deploy rights to ship code or reach production.

Secure code review

Secure code review is the examination of source code, usually a pull request diff, specifically to find security flaws such as missing authorization checks, injection, unsafe deserialization and leaked secrets, by a person, a static analysis tool, an AI reviewer, or a combination, before the change reaches production.

Attack surface

An attack surface is every point where an attacker can try to get into a system, affect it, or pull data out of it: internet-facing hosts, APIs, login pages, cloud services, internal network services, physical access and the people who can be tricked.