How Aevral reviews a PR
The mechanics of one Aevral pull-request review: what code it reads and never reads, how files are packed, which passes run, the grounding gate, verdicts, and the hard limits as numbers.
This page answers the question an engineer asks before trusting a review bot: what does it actually read, and what has to be true before it is allowed to say anything. The companion page What a PR review looks like covers the product behavior: triggers, plans, and what lands on your pull request.
The short answer
- Aevral is not SAST. It is an AI code review with mechanical gates on top, not a rules engine.
- On each push it reads the diff of every changed file, plus the full content at the head commit of the security-critical changed files, such as auth, identity, session, crypto, input handling, SQL, and routing files. It never reads git history, never runs a whole-repo retrieval for a review, never executes your code.
- It does not grep. There are no pattern rules. Three specialist passes read the same packed content: one for authorization, IDOR, and business logic; one for injection classes such as SQL, command, XSS, SSRF, path traversal, and unsafe deserialization; one for identity and AI-integration boundaries such as token and session flaws, prompt injection into tool calls, and unsafe model output.
- A finding is only admitted when the model can supply the full chain: the attacker, the input they control, the path to the sensitive operation, the missing control, the consequence, the evidence, and a fix. A parser rejects filler answers.
- The finding must then ground: the code snippet it cites must match, verbatim and uniquely, lines the pull request actually added (or, for a deleted check, lines it removed). An ungrounded finding is discarded mechanically, not by judgment.
- Only a "confirmed" verdict publishes, meaning proven by the diff itself. Findings the model could not prove are listed in the Check as unverified, never posted as comments. At most five findings publish per review, and the review says when more were held back.
- Silence is the default. A clean pull request gets one advisory Check and no comment noise. A new review notification is posted only for new findings at medium severity or above; inline comments only at high or above.
- It is probabilistic, not exhaustive. On our planted-vulnerability corpus of real multi-line bugs, the paid-plan lane publishes a median of 21 of 24 plants per run with zero published false positives across 72 clean-control runs; the free-tier lane published medians of 16 to 20 of 24 across same-day builds, with one false positive in the worst repeat. Misses happen and are published on the receipts page.
- It covers access control, business logic, injection, and identity classes only. No secrets, no dependencies, no security headers, no memory corruption, no denial of service, no code style.
- So it is complementary to rule-based SAST: keep your rules for exhaustive, deterministic, cheap coverage; add Aevral for cross-file authorization and business-logic reasoning that pattern rules cannot express.
What the review reads
- The diff of every changed file of the pull request.
- The full file content at the head commit for changed files whose paths look security-critical: tier one is auth, identity, sessions, permissions, crypto and secrets; tier two is input handling, SQL and data access, routing and API surface. At most 40 files, at most 40,000 characters each, and only when the packing has room next to their diff.
- The pull request title, its first 3,000 characters of description, and a bounded map of the changed files' trust boundaries.
- Nothing else. No git history. No whole-repo file listing. No execution. After a re-push, a GitHub compare call is used only to pick which of the pull request's own files changed since the last reviewed head; a force-push or a rebase triggers a full review again.
Generated files, lockfiles, translation data, test data, and eval outputs are skipped by rule and named in the Check, so you can see what was left out. A file whose path looks like it holds real secrets is never sent to the model at all; it is named in the Check as not sent. The exclusions cannot be used to dodge a review: a security-critical file is never excluded, and a pull request whose every file would be excluded is reviewed whole.
How files are packed
Files are read most security-relevant first: auth and identity before input and SQL, before configuration and CI, before other source, before documentation. Tests and fixtures always come after production code. The order is deterministic: the same file list always packs the same way.
The content is packed into model calls of at most 80,000 characters each. The diff of every changed file is packed first; the full head content of security-critical files joins where it fits. A single diff larger than a whole batch is cut at the batch size and the file is named as truncated. Files that fit in no batch are named as not read, with the reason.
The number of batches per review is bounded (the default is four, by plan), so a very large pull request gets its most security-relevant files reviewed and the Check states exactly what was covered. Every reviewable file ends the run with a disclosed fate: read, truncated, not read with a reason, excluded by rule, or withheld as sensitive.
The passes
Each packed batch is read by up to three passes:
- Authorization always runs: broken authorization, IDOR and BOLA, confused deputy, privilege propagation, business-logic abuse (replay, duplicate redemption, quota or billing bypass, ordering), cross-system identity errors, fail-open behavior, and removed security controls.
- Injection runs in the general scope: SQL and NoSQL injection, command injection, template injection, XSS, SSRF, path traversal, and unsafe deserialization.
- Identity and AI boundaries runs in the general scope: disabled TLS verification, non-cryptographic randomness in security tokens, token and session validation flaws, prompt injection that reaches a tool invocation, and model output consumed as code or HTML without validation.
Paid plans run on GLM-5.3 at low reasoning effort; the free tier runs on GLM-5.3 Flash. Both go through OpenRouter with a pinned provider list and a zero-retention requirement; see Security and data for where your code is processed.
The grounding gate
The model is told to quote 1 to 6 consecutive lines the pull request added, verbatim, including the vulnerable line. The gate then checks the quote mechanically against the diff:
- The quote (at least 16 normalized characters) must appear in one run of added lines, exactly once in the file, within 10 lines of the line the model cited.
- The published comment is re-anchored onto the line that actually matched.
- An ambiguous or distant match fails closed: no comment.
- A quote that only spans unchanged context lines can ground at best as unverified, never as confirmed. Pre-existing code beside a touched line can never publish as a finding of this pull request.
- For a security check the pull request deleted, the same gate runs against the removed lines.
If the model invents, paraphrases, or misplaces a snippet, the finding is discarded automatically. A confirmed finding that fails grounding is still counted internally, and the run then never posts a public clean verdict: a possible real bug we could not verify is not treated as evidence of a clean pull request.
Verdicts, the cap, and the unverified tier
Every candidate finding carries a verdict:
- Confirmed means the diff itself proves the path: the attacker, the input, the route to the sink, and the control that is missing or removed. An unseen compensating control is not a reason to downgrade; a control shown in the diff is. Only confirmed publishes.
- Probable means the finding passed the same grounding gate but its harm depends on a fact outside the diff. It is listed in the Check under "Unverified, worth a look", at most five, never as an inline comment, and never counted as a clean result.
- Hypothesis and dismissed findings never reach you. The model is instructed to attack its own findings before emitting them and to default to silence under uncertainty; silence on an uncertain pull request is a correct outcome.
At most five confirmed findings publish per review, most severe first, and the Check says how many more were held back. Duplicate findings of one root cause are merged mechanically: the same vulnerable line is one finding, whatever span each pass quoted, and the higher severity wins.
A worked finding
An abridged, paraphrased example from our planted-vulnerability corpus (the corpus is private; this is a planted test case, not a customer finding).
The pull request adds an optional parameter to a disconnect endpoint:
When integrationId is supplied, look up the integration by that id alone
and hard-delete it. No organization_id filter is applied on this path.The published finding, with its fields shortened:
- Attacker model: an authenticated owner of any organization that has a Slack integration.
- Attacker input: the
integrationIdparameter, which they control. - Path trace: the disconnect handler reads the integration by id alone, then deletes it and revokes its bot token.
- Missing control: the lookup is not bound to the caller's organization.
- Consequence: the attacker disconnects another organization's integration and forces revocation of its bot token.
- Verdict: confirmed; the added lines themselves drop the organization binding.
- Published as: one inline comment pinned to the added line that performs the unscoped lookup, with a suggested fix prompt for your coding agent.
The snippet quoted in the finding matched the added lines verbatim, so the comment posted. Had the model paraphrased the line, the finding would have been discarded, even though the bug was real. That is the trade the grounding gate makes.
What the Check says, honestly
The Check is advisory and never blocks a merge. The green clean sentence ("no findings across its full code-security scope") appears only when the run covered every reviewable file, no pass was skipped, nothing was withheld as sensitive, and no cost limit fired. A partial run says exactly that: "Partial coverage is not a clean bill of health", with the list of files not read and why. A review that hit its cost safety limit says so and tells you to review manually.
On a re-push, the summary review is edited in place; a new review (which sends GitHub a notification) is posted only when there are new findings, and the same finding is never announced twice.
The limits, as numbers
| Limit | Value |
|---|---|
| Pull requests larger than this are not reviewed at all (neutral Check) | 600 changed files, or 3,000,000 total diff characters |
| Model calls per review | up to 8 batches of 80,000 characters (default 4, by plan) |
| Full head content | up to 40 security-critical files, 40,000 characters each |
| Published findings per review | 5 confirmed, most severe first, the cut disclosed |
| Unverified items listed | 5 |
| Grounding window | 1 to 6 consecutive added or removed lines, within 10 lines of the citation |
| New review notifications | only for new findings at medium severity or above |
| Inline comments | only at high severity or above |
Is it SAST?
No. A rule-based SAST engine such as Semgrep matches your code against rules and dataflow patterns, deterministically, across the whole tree, in seconds. Aevral reads the diff like a security engineer would, reasons across files about who can reach what, and must prove each claim on an added line before it speaks. What rules cannot express (cross-file authorization, tenant scoping, business logic, a removed check) is what Aevral is for; what rules do best (exhaustive pattern coverage, secrets, dependencies, speed, determinism) stays with your SAST. Keep both. The comparison in detail: aevral.com/is-aevral-sast.
Measured behavior, misses included, is on the receipts page. What Aevral does not cover at all: Supported and not supported.