Your app works locally and breaks in staging. The two environments differ in eight settings, and none of them looks wrong on its own. EnvCause takes the working config, the broken config, and a command that shows the failure, then tells you which of those eight changes are needed to break it. In the bundled demo the answer is two.
$ envcause --good good.env --bad bad.env -- python examples/demo_app.py
Original differing variables : 8
Failure-inducing variables : 2
1-minimal failure-inducing change set:
FEATURE_NEW_AUTH: false -> true
JWT_ALGORITHM: HS256 -> RS256Why a developer would use it
You could diff the two files, but a diff tells you what changed, not which change matters. You could flip settings one at a time, but that only finds failures caused by a single setting. The failure in the demo needs both the new auth flag and the new signing algorithm. Each is fine alone, so one-at-a-time testing finds nothing.
EnvCause automates the search that you would otherwise do by hand, badly, at the end of a long day. It is useful when:
- a staging or CI deploy fails after a batch of environment changes and nobody knows which one did it;
- a config migration touched dozens of keys in YAML or TOML and you need the one or two that matter;
- you want to hand a teammate or an issue tracker a small reproduction instead of two 200-line files;
- a CI job fails on a config drift and you want the diagnosis attached to the run, not reproduced later on someone's laptop.
It is not a config validator. A schema checker tells you a value is invalid. EnvCause tells you which valid-looking combination breaks your application.
How it works
Git bisect narrows a history to one commit. EnvCause applies the same idea to a set of changes, using the classic ddmin delta-debugging algorithm. It starts from the good config and treats each difference as a switch that can be flipped to its bad value. It splits the switches into chunks, runs your command with some of them flipped, and keeps whichever subset still fails. It then splits finer until no single remaining change can be dropped.
The result is 1-minimal: remove any one change left in the set and the failure goes away. It is not guaranteed to be the globally smallest set. That is a deliberate trade, because chasing the true minimum would cost far more command runs, and 1-minimal is what a person needs in order to act.
Variables that exist on only one side work naturally. A key only in the bad file is a candidate that sets it, and a key only in the good file is a candidate that unsets it. For JSON, YAML, and TOML, changes are addressed by JSON Pointer path and objects are reduced recursively. Arrays are treated as single changes so every candidate stays a valid document.
Deciding what counts as a failure
By default, a non-zero exit code means the failure reproduced. That is often too loose, because a command can fail for unrelated reasons partway through a reduction and send the search down the wrong path. So the failure condition is configurable: a literal string or regular expression in the output, or a <failure> or <error> element in a fresh JUnit XML report.
The JUnit matcher deletes the previous report before every run, so a command that never writes a fresh report cannot be judged on a stale one.
For flaky failures, --repeat requires a candidate to fail on every run, and --verify-repeat lets the two baselines be checked more times than the search itself. A single lucky pass on the bad baseline should not send the whole search off course.
Running it where the failure lives
The same reduction can run against three targets. Locally is the default. With --docker-image, each candidate gets a fresh container, so candidates cannot contaminate each other. With --kube-pod, candidates run in an existing pod through kubectl exec, which suits failures that only happen with the real cluster network and mounts.
A reduction usually revisits the same subsets, so results are cached during the run. A persistent cache file stores SHA-256 fingerprints and pass/fail outcomes, never raw values, keyed by the inputs, environment, command, matcher, and repeat count. Independent candidates can also run in parallel where they do not share state. EnvCause rejects --parallel for structured files, JUnit reports, and pods rather than let candidates overwrite each other.
There is also a composite GitHub Action that runs the reduction in CI, writes a redacted JSON report, caches candidates between runs, and adds a result table to the job summary. envcause explain turns a saved report into a terminal diagnosis or a Markdown file for an issue or incident write-up, without re-running anything.
Handling configuration safely
Config files contain secrets, so the tool assumes so. It has no telemetry and no network client of its own. It redacts values whose names or paths look sensitive, such as TOKEN, PASSWORD, or KEY, in terminal output and reports, and writes real values only to the candidate and reproduction files you ask for, with owner-only permissions on POSIX systems. If a reduction fails or is interrupted, it restores a pre-existing candidate file.
One limit belongs to you, not the tool: EnvCause runs your command, often many times, with whatever access that command has. The README recommends a disposable staging pod or a fresh container, and sanitized copies of any production config.
My ownership
I designed the command-line interface, wrote the reduction and adapters in Python, built the structured-config layer, and wrote the test suite (42 tests). I also set up CI, the PyPI release, the GitHub Action, and the documentation. It is MIT-licensed, with no dependencies beyond PyYAML and tomli-w for structured formats.
Outcome
EnvCause is on PyPI (pipx install envcause, Python 3.10+) and in the GitHub Marketplace. It grew over four minor releases from a dotenv-only tool into one that handles structured files, flaky failures, Docker, Kubernetes, and CI. The latest patch release, 0.4.1, fixed stale JUnit reports, non-ASCII dotenv values, and file permissions for secret-bearing outputs. I have no adoption numbers to report.