Skip to content
Ayush Gupta

· 4 min read

Ten ways AI coding agents fake a green build (and how to catch them in the pull request)

By Ayush Gupta, AI Engineer in Bengaluru

When a coding agent can't fix a failing build, it sometimes makes the build pass anyway.

Not by fixing the bug. By changing the test, the check, or the pipeline until the red goes away. The build turns green, the reviewer sees a passing check, and the change gets merged.

I built Greenwash, an open-source GitHub App, to catch exactly this. This post covers the ten patterns it looks for, why a general AI code review is the wrong tool for the job, and the design decisions that make the signal useful instead of noisy.

What it looks like

Here are two hunks from a pull request where an agent was asked to fix a failing checkout test:

# tests/test_checkout.py
- assert result.total == 100
+ assert total > 0
# src/pricing.py
+ if user_id == 4821:
+     return expected_result

The first change weakens an assertion until it can no longer fail. The second hard-codes the answer the test expects for one specific input. Both make CI green. Neither fixes anything.

In a small diff you would spot these. In a 40-file pull request, written quickly by an agent and reviewed quickly by a person, they are easy to miss. That is the whole problem: these changes are small, and they hide well.

Here is a short walkthrough of Greenwash catching changes like these:

The ten patterns

  1. Weakened test assertion. An exact check becomes a loose one.
  2. Test skipped or disabled. A skip marker, a commented-out test, an early return.
  3. Expected value rewritten without any stated change in behavior.
  4. Hard-coded answer for specific inputs, like the user_id == 4821 branch above.
  5. Logic replaced with a placeholder. The real implementation becomes a stub that satisfies the test.
  6. Error silently swallowed. A try with an empty except, so the failure disappears.
  7. Type or lint check suppressed. An ignore comment instead of a fix.
  8. CI step removed or made non-blocking.
  9. Test file deleted.
  10. Snapshot updated without any source change.

None of these are always wrong. Sometimes you really do mean to update an expected value or skip a flaky test. That matters for the design, and I'll come back to it.

Why a general AI review doesn't solve it

The obvious approach is to point an LLM at the diff and ask "is anything wrong here?". In practice that tends to produce an essay: a long list of observations of mixed importance, with no clear signal about which line to look at.

Greenwash does the opposite. It reads every changed hunk of a pull request and asks narrow yes/no questions about it, such as:

Does this delete a test assertion or replace it with a weaker one?

Each question is answered with a probability by TypeSafe's structured-judgment models. Eight of the ten checks work this way. The other two, like a deleted test file, are plain code rules, because you don't need a model to notice a file is gone.

The principle I followed: the model handles judgment, code handles policy. The model answers "how likely is it that this hunk weakens an assertion?". Code decides the thresholds, the severity, and what gets reported.

Design decisions, and what each one costs

Probabilities, not verdicts. A flagged change may be intentional, so every finding shows how confident it is. A finding is flagged at 0.8 by default, and anything from 0.5 up is marked as worth a second look. The cost: you need thresholds, and they are a judgment call.

Comment mode by default. Greenwash reports through a check run with line annotations and a single pull request comment that updates in place on every push. It never blocks a merge unless a repository opts in with mode: check. The cost: enforcement is opt-in.

Config is read from the base commit. Repositories tune Greenwash with .github/greenwash.yml, and the file is read from the pull request's base, not its head. Otherwise an agent could switch off the very check that would catch it. The cost: config changes only apply after they merge.

Fail neutral. If Greenwash itself errors, the check completes as neutral instead of failing your build. The cost: an outage means no findings for that run.

One request per hunk. Each request carries only the hunk's diff, the file path, and the pull request's title and description. No whole files, no database, nothing stored; findings live only on GitHub. The cost: each hunk is judged without the rest of the file.

You can also add your own checks as plain questions. For example, a team could flag any change to billing logic:

rules:
  - id: billing
    title: Billing logic changed
    question: "Does `hunk.diff` change how customers are charged, refunded, or invoiced?"
    applies_to: [source]
    severity: warning

How I evaluated it

The repository ships a labelled evaluation set. For every model-answered check there are two real examples of greenwashing and two legitimate look-alike changes: 32 cases in total. All 32 classify correctly at the default thresholds.

To be clear about what that means: the set is small and hand-written. It works as a regression suite that stops a change to a question or threshold from breaking a check. It is not a measure of real-world precision. Reports of false positives and misses are the most useful contribution anyone can make to the project.

Known limits

  • GitHub only for now, no GitLab or Bitbucket.
  • Each hunk is judged on its own, without the rest of the file or repository.
  • Very large pull requests are capped at 300 hunks, and each hunk at 8,000 characters.
  • Every finding is a probability. Intentional changes can be flagged, and subtle cheating can be missed.

Try it

Greenwash is open source under Apache-2.0 and self-hosted: you register your own GitHub App and deploy with Docker or Railway. The setup guide is in the repository.

If you run coding agents against real repositories, I'd like to hear which patterns you see that aren't on this list.