Skip to content

AI Code Review for SRE & Reliability Teams

Argus reviews a change with the code that depends on it in the prompt, and re-checks earlier findings and bug-labeled issues against diffs that touch their files.

Last updated

What problems do SRE & Reliability Teams face in code review?

  • Changes to shared infrastructure code break downstream services in ways no single PR review can predict
  • Incident post-mortems identify the same root causes repeatedly — but that knowledge doesn't flow back into the review process
  • On-call engineers reviewing PRs at 2am miss cross-module dependencies that cause the next page
  • Manual blast radius analysis is slow and incomplete — you can't trace every affected endpoint by hand

How does Argus fit your workflow?

  • Before review, Argus walks a code graph of the default branch two levels back from the changed files and gives the reviewer model up to 50 dependents, plus the source of up to 3 direct ones. The list is prompt context for one repo; it is not posted on the PR
  • Issues labeled argus or bug, and critical or warning findings from earlier reviews, are stored as failure scenarios; an issue's files are the paths named in its body, so an issue that names none is never re-checked. When a PR touches their files, an LLM re-checks up to 5 against the diff and the review lists the ones it confidently judges still broken. Nothing is executed
  • Review memory puts past findings, dismissed false positives, and team-written rules into review prompts. Argus does not read post-mortems; an incident reaches it as an issue labeled argus or bug, a scenario added through the API, or a pattern someone teaches with @argus-eye remember (the handle is your GitHub App's slug; argus-eye is the default)
  • A changed file that 5 or more other files depend on is flagged as a choke point in its review prompt, and one with 3 or more past critical or warning findings as a hotspot. Argus does not compute a blast-radius score

Which Argus features matter most for your team?

Earlier-finding re-checks
Stored scenarios for the touched files are re-checked by an LLM against the diff text, and confident 'still broken' verdicts appear on the PR as earlier findings that may still apply. No code runs, so failures that only show up at runtime are out of reach
Dependents in the review prompt
The reviewer model sees files that call or import the changed code, from a graph refreshed on each push or merge to the default branch. Dependents in other repos are not traced
Review memory
Findings, patterns, and team feedback from earlier reviews feed later ones, and a finding that closely matches a stored pattern cites it, for example 'Matches a prior fix in PR #N'
Choke-point and hotspot flags
Files many others depend on, or that have produced findings before, get an explicit warning in their review prompt. The dashboard's Memory → Files view lists each file's fan-in

By default, each review re-checks up to 5 stored failure scenarios, one LLM call each, against diffs truncated to 2,000 characters per file; no code is executed

— Argus source, backend/internal/pipeline/simulation.go

Run Argus on your own repositories

Open source under AGPL-3.0, with no paid tier and no feature gating.

Self-hosted only: Docker Compose or Fly.io, Postgres with pgvector, a GitHub App and a Clerk app you create, your model keys and an embeddings endpoint.