Semantic lint bench average precision on 214 public pairs: Jev 95.0%, DiffusionGemma Jev 87.0%, Span-01 79.6%, CLM-v0.1-8B 50.5%, and SemiF 50.0%.

The semantic lint bench simulates a developer writing a short ensure rule in perch.yaml to describe correct behavior. We checked the same rule against a method before and after its fix. These examples come from the public pairs:

Bug Lint rule
SwiftNIO: compressed bytes can remain buffered for an empty response in PartialHTTPResponse.flush Finalize compression on completed responses even with an empty input body, so buffered bytes are emitted.
tcpdump: a protocol-header read can exceed packet bounds in wb_prep Check both declared packet length and captured-buffer bounds before reading a protocol header.
Jenkins Kubernetes plugin: doFillCloudItems can expose cloud names without administrator permission Cloud-listing endpoints must check administrator permission before returning configured cloud names.

Jev scored the broken version as more likely to violate its rule in 206 of 214 public pairs (96.3%). That is strong semantic linting performance on a stated requirement. A behavior rule can serve as a regression check: assert what the code must do, then run the same check after later edits.

Benchmark criteria

We assembled 2,718 before-and-after bug pairs from open-source fixes, largely drawn from SWE-Rebench and SWE-bench, and 1,750 security pairs linked largely to fixes in the GitHub Advisory Database. From these sets, we selected 285 pairs for the semantic lint bench.

For each method, Perch used its AST analysis to gather the source, imports, callers, callees, and available call-graph edges. Jev saw one version at a time, with that context but without the patch. The bug test asked whether the method had a reachable defect. The security test used a pack of 30 CWE-specific questions, filtered by language and drawn from MITRE’s 2025 Top 25 and related weaknesses.

Results

Task Public pairs Jev AP Precision Recall Run cost
Find a bug 2,220 46.2% 55.6% 5.9% $0.620
Find a security issue 1,311 51.5% 53.1% 34.6% $0.507
Semantic lint bench 214 95.0% 99.1% 50.0% $0.048

The broad bug scan found 109 of 1,861 labelled buggy methods; the security scan found 395 of 1,142 vulnerable methods. On the semantic lint bench, Jev flagged 108 methods at the 70% cutoff. 107 were broken, and it found 107 of the 214 broken methods.

AP compares scores across cutoffs. Precision and recall use Perch’s configured alert cutoffs (70% for semantic lints). Costs cover these public benchmark runs.

Jev versus four other decision models

We also scored the same methods with four other decision models. Higher average precision means issue-bearing methods tend to rank ahead of fixed and benign methods across score cutoffs.

Model Bugs Security Semantic lint bench Public run cost Context window
Jev 46.2% 51.5% 95.0% $1.175 64k
DiffusionGemma Jev 46.8% 48.2% 87.0% $0.465 64k
Span-01 43.8% 47.6% 79.6% $0.447 22.9k
SemiF 44.6% 43.9% 50.0% $0.465 256k
CLM-v0.1-8B 42.1% 42.9% 50.5% ~$1.652 6k

DiffusionGemma Jev was the strongest open-weight competitor in these runs. It narrowly led Jev on broad bug ranking (46.8% versus 46.2% average precision), while Jev led on security (51.5% versus 48.2%) and the semantic lint bench (95.0% versus 87.0%).

SemiF and CLM marked almost every security method as vulnerable. Both reached 100% recall, but generated about 1,480 false alerts each on the public security set.

The private, repository-disjoint semantic lint pairs told a similar story: Jev reached 94.3% average precision on 71 pairs, against 95.0% on the public 214. The full results include private-set rankings, precision, recall, F1, alert counts, costs, and public per-method predictions.

Data and reproduction

The bug pairs and security pairs cover Python, Rust, Go, JavaScript, TypeScript, Java, C#, C, PHP, and other languages. Public and private splits have no source repositories in common.

The public semantic lint bench pairs, benchmark runner, and frozen question pack are available on Hugging Face.

Jev ran through TypeSafe’s API; DiffusionGemma Jev and SemiF ran on Beam, CLM on a Runpod H100, and Span-01 through Respan’s API. The recorded API charges or GPU time give the costs above. The public pair files, runner, and predictions are linked from the results page.

Try Systemone models with Perch

When triaging or fixing a bug, ask your coding agent to turn the behavior the code must preserve into a persistent ensure rule in perch.yaml. The Perch agent skill shows how to write the rule and check it against code that should pass and code that should fail. Run perch scan --filter rule=<name> after an edit. In CI, perch scan --since origin/main checks changed code alongside tests.