# Can Jev find bugs and security vulnerabilities?

Benchmarking Jev and other open source models on code fixes and semantic lint rules

Source: https://perchscan.com/blog/benchmarking-behavior-driven-code-checks/
Published: 2026-09-28

<picture>
  <source media="(max-width: 600px)" srcset="/blog/benchmarking-behavior-driven-code-checks/header-mobile.png" />
  <img src="/blog/benchmarking-behavior-driven-code-checks/header.png" width="2000" height="800" alt="Semantic lint bench average precision on 214 public pairs: Jev 95.0%, DiffusionGemma Jev 87.0%, Span-01 79.6%, CLM-v0.1-8B 50.5%, and SemiF 50.0%." />
</picture>

The semantic lint bench simulates a developer writing a short `ensure` rule in `perch.yaml` to describe correct behavior. We checked the same rule against a method before and after its fix. These examples come from the [public pairs](https://huggingface.co/datasets/perchscan/perch-lint-bench-public):

| Bug | Lint rule |
| --- | --- |
| SwiftNIO: compressed bytes can remain buffered for an empty response in [`PartialHTTPResponse.flush`](https://github.com/apple/swift-nio-extras/blob/6740bf98c2b758dd4511a6d0076fbaff8f1e0a82/Sources/NIOHTTPCompression/HTTPResponseCompressor.swift) | Finalize compression on completed responses even with an empty input body, so buffered bytes are emitted. |
| tcpdump: a protocol-header read can exceed packet bounds in [`wb_prep`](https://github.com/the-tcpdump-group/tcpdump/blob/13ab8d18617d616c7d343530f8a842e7143fb5cc/print-wb.c) | Check both declared packet length and captured-buffer bounds before reading a protocol header. |
| Jenkins Kubernetes plugin: [`doFillCloudItems`](https://github.com/jenkinsci/kubernetes-plugin/blob/203cde6529ced26944021c77ee27991e3578976c/src/main/java/org/csanchez/jenkins/plugins/kubernetes/pipeline/PodTemplateStep.java) can expose cloud names without administrator permission | Cloud-listing endpoints must check administrator permission before returning configured cloud names. |

Jev scored the broken version as more likely to violate its rule in **206 of 214 public pairs (96.3%)**. That is strong semantic linting performance on a stated requirement. A behavior rule can serve as a regression check: assert what the code must do, then run the same check after later edits.

## Benchmark criteria

We assembled **2,718 before-and-after bug pairs** from open-source fixes, largely drawn from [SWE-Rebench](https://huggingface.co/datasets/nebius/SWE-rebench) and [SWE-bench](https://github.com/SWE-bench/SWE-bench), and **1,750 security pairs** linked largely to fixes in the [GitHub Advisory Database](https://github.com/advisories). From these sets, we selected **285 pairs** for the semantic lint bench.

For each method, [Perch](https://github.com/lakeday-org/perch) used its AST analysis to gather the source, imports, callers, callees, and available call-graph edges. [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) saw one version at a time, with that context but without the patch. The bug test asked whether the method had a reachable defect. The security test used a pack of 30 CWE-specific questions, filtered by language and drawn from [MITRE's 2025 Top 25](https://cwe.mitre.org/top25/archive/2025/2025_cwe_top25.html) and related weaknesses.

## Results

| Task | Public pairs | Jev AP | Precision | Recall | Run cost |
| --- | ---: | ---: | ---: | ---: | ---: |
| Find a bug | 2,220 | 46.2% | 55.6% | 5.9% | $0.620 |
| Find a security issue | 1,311 | 51.5% | 53.1% | 34.6% | $0.507 |
| Semantic lint bench | 214 | **95.0%** | **99.1%** | 50.0% | $0.048 |

The broad bug scan found 109 of 1,861 labelled buggy methods; the security scan found 395 of 1,142 vulnerable methods. On the semantic lint bench, Jev flagged 108 methods at the 70% cutoff. **107 were broken**, and it found 107 of the 214 broken methods.

*AP compares scores across cutoffs. Precision and recall use Perch's configured alert cutoffs (70% for semantic lints). Costs cover these public benchmark runs.*

## Jev versus four other decision models

We also scored the same methods with four other decision models. Higher average precision means issue-bearing methods tend to rank ahead of fixed and benign methods across score cutoffs.

| Model | Bugs | Security | Semantic lint bench | Public run cost | Context window |
| --- | ---: | ---: | ---: | ---: | ---: |
| [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) | 46.2% | **51.5%** | **95.0%** | $1.175 | [64k](https://docs.typesafe.ai/models) |
| [DiffusionGemma Jev](https://github.com/razorback16/openjev) | **46.8%** | 48.2% | 87.0% | $0.465 | [64k](https://codiv.ai/docs/models) |
| [Span-01](https://www.respan.ai/docs/documentation/span-01/concept) | 43.8% | 47.6% | 79.6% | $0.447 | [22.9k](https://huggingface.co/datasets/perchscan/benchmark-results/blob/main/predictions/current/security/span01.jsonl) |
| [SemiF](https://github.com/TheoLeeCJ/SemIf) | 44.6% | 43.9% | 50.0% | $0.465 | [256k](https://huggingface.co/Qwen/Qwen3.5-4B) |
| [CLM-v0.1-8B](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B) | 42.1% | 42.9% | 50.5% | ~$1.652 | [6k](https://huggingface.co/datasets/perchscan/benchmark-results/blob/main/adapters/clm.py) |

DiffusionGemma Jev was the strongest open-weight competitor in these runs. It narrowly led Jev on broad bug ranking (46.8% versus 46.2% average precision), while Jev led on security (51.5% versus 48.2%) and the semantic lint bench (95.0% versus 87.0%).

SemiF and CLM marked almost every security method as vulnerable. Both reached 100% recall, but generated about 1,480 false alerts each on the public security set.

The private, repository-disjoint semantic lint pairs told a similar story: Jev reached **94.3%** average precision on 71 pairs, against 95.0% on the public 214. The [full results](https://huggingface.co/datasets/perchscan/benchmark-results) include private-set rankings, precision, recall, F1, alert counts, costs, and public per-method predictions.

## Data and reproduction

The [bug pairs](https://huggingface.co/datasets/perchscan/bug-pairs-public) and [security pairs](https://huggingface.co/datasets/perchscan/security-pairs-public) cover Python, Rust, Go, JavaScript, TypeScript, Java, C#, C, PHP, and other languages. Public and private splits have no source repositories in common.

The public [semantic lint bench pairs](https://huggingface.co/datasets/perchscan/perch-lint-bench-public), [benchmark runner, and frozen question pack](https://huggingface.co/datasets/perchscan/benchmark-results/tree/main) are available on Hugging Face.

Jev ran through TypeSafe's API; DiffusionGemma Jev and SemiF ran on Beam, CLM on a Runpod H100, and Span-01 through Respan's API. The recorded API charges or GPU time give the costs above. The public pair files, runner, and predictions are linked from the [results page](https://huggingface.co/datasets/perchscan/benchmark-results).

## Try Systemone models with Perch

When triaging or fixing a bug, ask your coding agent to turn the behavior the code must preserve into a persistent `ensure` rule in `perch.yaml`. The [Perch agent skill](https://docs.perchscan.com/skill/) shows how to write the rule and check it against code that should pass and code that should fail. Run `perch scan --filter rule=<name>` after an edit. In CI, `perch scan --since origin/main` checks changed code alongside tests.
