My Postgres migration linter analyzed nothing for four and a half months

From 0.5.0 through 0.7.0, pgfence exited 0 after analyzing nothing when launched through a symlinked path, including npx, node_modules/.bin, pnpm, global installs, and its own GitHub Action. A user found it. Here is the bug and the fail-closed fix.

If you ran pgfence through npx, through node_modules/.bin/pgfence, through a global install, under pnpm, or through the official flvmnt/pgfence GitHub Action, at any version from 0.5.0 to 0.7.0, it printed nothing, analyzed nothing, and exited 0. --ci communicates through the exit code and nothing else, so the gate was green. Every migration in every one of those pull requests went through unchecked, and the check mark next to it said the opposite.

The guard responsible went in on 2026-04-15 as part of a release-hardening commit. 0.5.0 shipped two weeks later. Nobody noticed for four and a half months. 0.4.1 and earlier are fine, 0.8.0 fixes it, and this post is the mechanism, the blast radius, and the four things I changed so that the next version of this bug is loud instead of quiet.

The bug is one line

dist/index.js is simultaneously a bin entry and an export target, so it has to know whether it was launched or imported. The usual ESM way to ask that, before import.meta.main existed on most runtimes:

const __filename = fileURLToPath(import.meta.url);
const isMainModule = process.argv[1] != null && path.resolve(process.argv[1]) === __filename;

// Commander definitions are registered here.

if (isMainModule) {
  program.parse();
}

import.meta.url is the URL of the module Node actually loaded, and Node has already resolved it through realpath. path.resolve(process.argv[1]) never touches the filesystem. It makes the path absolute and normalizes . and .., lexically, and stops. When any component of the launcher path is a symlink, those two strings name the same file on disk and are not equal as strings.

npm and yarn classic install a bin by writing a symlink at node_modules/.bin/pgfence pointing at ../@flvmnt/pgfence/dist/index.js. So argv[1] was the .bin path, import.meta.url was the dist path, isMainModule was false, program.parse() never ran, and the module fell off its own end with nothing registered and nothing to do. That is an exit 0. A clean one. There is no error to catch because nothing failed.

Every launcher below resolves through a symlink, and every one of them analyzed nothing:

  • node_modules/.bin/pgfence under npm and yarn classic
  • npx @flvmnt/pgfence, because the npx cache entry is a symlink
  • a global npm install -g bin
  • pnpm’s node_modules/@flvmnt/pgfence package symlink
  • a Windows directory junction
  • the flvmnt/pgfence GitHub Action, which runs through npx

There is a seventh case with no symlink in it at all. node ./node_modules/@flvmnt/pgfence puts a directory in argv[1], and no comparison between a directory path and a file path can ever be true, so that failed identically on every runtime without import.meta.main.

Why it survived so long

Windows looked healthy, and that is the whole reason this lasted into September. npm on Windows does not write a symlink for a bin, it writes a .cmd shim, and the shim execs the real script path. argv[1] arrives already canonical, the two sides compare equal, and the CLI runs. The platform people expect to be broken was the only one that worked.

The other half is that development never touched the broken path. node dist/index.js and pnpm build && node dist/index.js both hand Node a real file. Every test ran the built artifact by real path, which is exactly the one invocation the bug cannot reach. I was testing the code and not the launcher.

Blast radius, with the part I cannot measure marked

I am not going to perform an outage I have no evidence of. pgfence had 10 stars on GitHub. npm reported roughly 13,000 downloads a month, and that number could not distinguish people from CI runners, Docker layer caches, and registry mirrors. It could not tell me how many pipelines were affected, or whether the count was closer to five or five hundred.

The one hard number came from the report itself: once the binary actually ran, 16 of the reporter’s 17 migrations produced findings. That is one repository, and it is what the gate would have said in each of those pull requests, one at a time.

The honest summary is the shape rather than the size. Every affected run was a run somebody had deliberately configured, on the surface where a safety tool is supposed to matter most, and every one returned a pass that meant nothing. The Action is the part of the list I like least. It was the surface people trusted most, and it was a check mark that never ran.

The person who found it

@CarbonDude, in issue #2. He filed it with the root cause already written out: path.resolve is lexical, import.meta.url has already been realpath-resolved by Node, so the two were never going to compare equal through a symlinked launcher. He also worked out why Windows was unaffected, which is the detail I would have chased longest on my own. There was nothing left to diagnose, only to fix. He is credited in the 0.8.0 changelog entry, which links back to the issue.

A safety gate that produces no output and exits 0 is indistinguishable from a pass and should never be trusted as one.

A missing linter and a silent linter are not the same failure. A team with no migration checks knows it has no migration checks and stays appropriately nervous about the ALTER TABLE in the diff. A team with a green check from a tool that ran zero files has been handed confidence it did not earn. The tool manufactured it. That is worse than absence, and it is why the fatal failure mode here was never false positives. It is false negatives and false safety implied.

What 0.8.0 shipped

All of it was written, built, and verified against the built binary through a symlinked launcher.

The guard is fixed in both bin targets. Both sides are canonicalized by the same function now, realpathSync.native where the platform has it, with import.meta.main as a fast path on runtimes that provide it and case-insensitive comparison on win32 and darwin. When argv[1] is a directory, the guard redoes the package.json main and index.js resolution Node itself did and compares that. Every branch either leaves the previous answer alone or turns a false into a true, so no launcher that worked in 0.7.0 can stop working in 0.8.0.

A contradiction check refuses to exit 0 in silence. If the process was launched under the name pgfence or pgfence-lsp and the guard still cannot confirm it is the entry module, it writes a diagnostic naming argv[1], import.meta.url, platform, and runtime, then exits 2. Whatever the next version of this bug looks like, it should be loud.

pgfence-lsp carried the identical defect. Under pnpm and through node_modules/.bin/pgfence-lsp, the language server started and exited immediately with nothing on stdout or stderr. Editors showed a server that died without an error, which reads to a user as “no problems found.” The same comparison caused it and the same fix covers it.

A run that analyzes zero SQL statements now exits 2. An empty file, a comment-only file, or an ORM migration whose up() contained no recognized query call used to print [SAFE], No dangerous statements detected., and Coverage: 100%, then exit 0. Zero of zero statements is not 100% coverage and it is not a pass. That was a second, independent way to get a green check out of a run that checked nothing, and it did not need a symlink.

One practical warning: an upgrade can make a pipeline fail, and fail large. That is accumulated migration risk becoming visible, not a new false positive. To see the whole picture in one run without blocking on it:

pgfence analyze --ci --max-risk critical --no-lock-timeout \
  --no-statement-timeout prisma/migrations/*/migration.sql

Then ratchet down. The two timeout flags are load bearing: --max-risk critical on its own still exits 1 because a missing SET lock_timeout is an error-severity policy violation, and those fail a --ci run regardless of --max-risk.

Telemetry, and the honest reason for it

0.8.0 also added anonymous usage counts. The justification is this bug rather than anything I would like to know about you. A release in which the CLI ran zero times through a symlinked bin would have shown up as zero events inside a week instead of inside four and a half months.

There is one question it exists to answer: how many real humans use this, as opposed to CI runners and registry mirrors, which is the one thing npm download counts cannot tell me. An event contains numbers, booleans, and values from closed lists: a random install ID, the command, the pgfence, Node, and OS versions, whether the run was in CI, how long it took, and counts of findings by severity. No SQL, file names, paths, table or column names, rule IDs, repository names, hostnames, usernames, environment variable values, database URLs, or IP addresses.

The first run on a machine prints a notice and sends nothing. The language server never sends anything. PGFENCE_TELEMETRY=0, DO_NOT_TRACK=1, pgfence telemetry disable, and telemetry = false in .pgfence.toml each turn it off on their own. PGFENCE_TELEMETRY_DEBUG=1 prints the exact payload to stderr without sending it. Every field and the limits of the resulting numbers are on the telemetry page.

One consequence of the bug went straight into the schema: a run that analyzes nothing is recorded as an errored run rather than a clean one, so an always-on gate that stops firing in the field shows up in the numbers instead of looking like a healthy analysis.

In the first 36 hours after telemetry went live, it recorded five distinct CI install IDs and zero interactive ones. That sample is too small to support a market conclusion. It was still the first real measurement this project had ever had.

The part a CLI cannot fix

Making pgfence run again does not make anyone keep it on. The CLI honors -- pgfence-ignore on the line before a statement and drops the finding before it reaches the result. That is right for a tool you run on your own machine and wrong for a centrally enforced gate, because the author of the migration is the person the gate exists to check. One comment line is otherwise a self-granted exemption with no reason, expiry, or audit trail.

That is the argument for pgfence Cloud, which runs the same analyzer as a GitHub App and refuses the directive, recording the attempt as a finding instead. Exemptions require a written reason and an expiry, and the run records who granted them. It is built and proven end to end against real pull requests, and it is not generally available. Today it is a design partner program, not something you can sign up for. The CLI stays free, and so does the required check itself, because a gate you have to pay to keep is a gate somebody will eventually turn off.

0.8.1 is on npm. If the upgrade turns up anything that looks wrong rather than merely newly red, open an issue.

← All posts