AI slop can look fine at first glance and pass all your tests. When you review the slop, you notice that it has some problem, like error handling that isn’t done correctly or docstrings that don’t match the code. You reject the PR and ask for it to be fixed. Write each kind of slop that you’re rejecting as a rule in your perch.yaml. Then, run Perch on your pull requests. When a PR contains a kind of slop that you’ve written a rule for, the build will fail before anyone has to review the PR.
$ perch scan --since main --filter type=lint
storage.py
ID Line Severity Type Confidence Problem Method
0f611fdb 27 - lint 89% docstrings-match-code storage.py
c5448501 27 - lint 79% no-hidden-errors storage.py
✖ 2 problems in 1 file, all failing
That’s a PR that contains a new function load_orders() with error handling that doesn’t match the docstring. The function swallows all errors and just skips files that can’t be loaded, but the docstring says that it will raise an error if a file can’t be loaded. In contrast, a PR with a correct load_orders() function passes the Perch check. Below, you can see the runs for both PRs and the rules that were used to check them.
AI slop code examples
This example project contains an order service with many functions, each of which has a different bug. Most of the bugs are kinds of AI slop. You can clone the example project and try it out yourself. Below is a table of how each file in the project scores against the three rules below. This was run with Perch 0.3.5.
| File | Slop in the source, for example | Perch result |
|---|---|---|
storage.py |
load_order() returns {} on any error, though its docstring says it raises |
no-hidden-errors 93%, docstrings-match-code 86% |
checkout.py |
can_fulfil() returns True when any item is in stock, not every item |
docstrings-match-code 92% |
cart.py |
cheapest() crashes on an empty cart instead of returning None |
docstrings-match-code 91% |
inventory.py |
restock() returns a count instead of the SKUs it touched |
docstrings-match-code 91%, no-placeholders 58% |
auth.py |
cancel_order() cancels any order, not only the customer’s own |
docstrings-match-code 83% |
Each rule applies to the entire file, so it is only reported once. The values in the middle column are based on reading the code in each file. Many files have multiple kinds of slop. These results were obtained without a confidence floor, which is why a borderline 58% shows up. You’ll set a confidence floor below.
Write the slop you keep rejecting as rules
Once you’ve set up Perch, you can start writing rules to detect kinds of slop that you want to avoid. Each kind of slop should be a separate rule. Use perch rules add with the --each file flag to write a rule as a question about each file in your repository. Adding these three rules created a file called perch.yaml with the following contents:
- name: no-hidden-errors
each: file
where: "**/*.py"
ensure: >-
Functions let errors reach the caller instead of catching them and returning
empty data.
- name: docstrings-match-code
each: file
where: "**/*.py"
ensure: Every function does what its docstring says.
- name: no-placeholders
each: file
where: "**/*.py"
ensure: >-
No function is a stub or placeholder that reports success without doing the
work.
You can now use Perch to check if a file contains any of the kinds of slop that you’ve written rules for. Just run perch check on the file you want to check.
$ perch check storage.py
storage.py:1
3 checks, 2 broken.
Confidence Rule Description
93% no-hidden-errors Functions let errors reach the caller instead of catching them and
returning empty data.
86% docstrings-match-code Every function does what its docstring says.
In this case, perch check found two kinds of slop in the file. After you fix the file, perch check will report that there are no issues and return an exit code of 0. In the above example, it would return an exit code of 3. Make sure to change the where clause to match the languages that you’re using.
Fail the pull request in CI
When you add a rule using perch rules add, it doesn’t have a confidence floor, which means that the PR with the correct load_orders() function wouldn’t pass. The issue is that Perch was only 55% confident that the docstrings don’t match the code, which is barely over 50%, so it’s a weak signal. To fix this, you can set a confidence floor for each rule. Here, each rule has been given a confidence floor of 70%.
$ perch rules edit no-hidden-errors --min 70
Changed no-hidden-errors.
$ perch rules edit docstrings-match-code --min 70
Changed docstrings-match-code.
$ perch rules edit no-placeholders --min 70
Changed no-placeholders.
Now, when you run perch scan on the PR that contains the illustrative agent-written load_orders() function, it will report that there are two kinds of slop.
def load_orders(directory):
"""Load every order in a directory. Raises if any file can't be read."""
orders = []
for name in sorted(os.listdir(directory)):
try:
orders.append(load_order(os.path.join(directory, name)))
except Exception:
pass
return orders
In contrast, a PR with a correct load_orders() function will pass the Perch check. The function allows errors from load_order() to bubble up, which is correct.
def load_orders(directory):
"""Load every order in a directory. Raises if any file can't be read."""
return [load_order(os.path.join(directory, name)) for name in sorted(os.listdir(directory))]
$ perch scan --since main --filter type=lint
✓ nothing to report
When run on the PR with the slop function, perch scan would return an exit code of 3, which would cause the CI job to fail. When run on the PR with the correct function, perch scan would return an exit code of 0, which would allow the CI job to pass. In CI, you should make sure to fetch the base branch before running perch scan, then use the --since flag to only check the files that have changed since the base branch. See the GitHub Actions guide for an example of how to do this.
Stop the slop at the source in Claude Code
If you’re using a coding agent that supports hooks, like Claude Code, you can run Perch at the end of each turn to prevent the agent from returning slop your rules describe. Claude Code has a hook called Stop that is run whenever Claude is about to end a turn. If the hook exits with a non-zero exit code of 2, then Claude will stay in the current turn and read the output of the hook from stderr. Perch exits with a non-zero exit code of 3 whenever a rule is broken, so you can write a simple script to convert the exit code to 2. Save the script in .claude/hooks/perch.sh, and make it executable with chmod +x .claude/hooks/perch.sh.
#!/bin/sh
# Claude Code runs this when it is about to finish. It checks each file Claude changed
# against perch.yaml. perch check exits 3 when a rule is broken; exiting 2 keeps Claude
# working and shows it the findings.
cd "$CLAUDE_PROJECT_DIR" || exit 1
status=0
IFS='
'
for file in $(git diff --name-only HEAD; git ls-files --others --exclude-standard); do
[ -f "$file" ] || continue
perch check "$file" >&2
[ $? -ne 3 ] || status=2
done
exit $status
Then, add the hook to your .claude/settings.json file.
{
"hooks": {
"Stop": [
{ "hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/perch.sh" }] }
]
}
}
The recorded session used Claude Code 2.1.284. The example project was checked out, and the rules were committed. The prompt asked Claude to write a new function save_order() in storage.py. Claude wrote a clean function and tried to end the turn, but the hook kept it in the turn and told it about the existing function load_order(). This happened twice, with confidences of 88% and 80% the first time, and 87% and 81% the second time. On the next attempt to end the turn, Claude removed the try/except that was swallowing the error, and the turn ended. The entire session took 10 turns. The session was recorded before the confidence floor was set, but all four findings had confidences above 70%, so the results would be the same with a floor.
Score variation and limits
Note that the scores can vary slightly each time you run Perch, since it’s asking the model a question each time. For example, the PR with the slop function had scores of 86% and 78% when there was no confidence floor, but when the PR was run at the top of this page, the scores were 89% and 79%. Perch is not designed to detect whether a file was written by an AI. It is designed to detect whether a file follows the rules that you’ve written. This means that it can be used to detect slop in files that were written by humans as well. It also means that Perch can only detect the kinds of slop that you’ve written rules for. In the Claude Code session above, Claude pointed out that the function append_line() still doesn’t close the file it opens, but there wasn’t a rule for that. Perch is a complement to tests, not a replacement. You should still write tests for your code, but if you find that your tests aren’t catching some kinds of slop, you can write rules for those kinds of slop and use Perch to catch them.