Skip to content

Argus documentation

Setup, the review pipeline, memory, configuration and bot commands.

Last updated

Getting started

Argus runs on your own infrastructure. Five steps to your first review, and one to tune it. The full reference is docs/self-hosting.md in the repo.

  1. Create your GitHub App

    In GitHub Developer settings, create an App under the org or personal account whose repos Argus will review, with its webhook at https://<your-backend-host>/webhooks/github and a webhook secret. Set the Callback URL to https://<your-dashboard>/github/callback, check Request user authorization (OAuth) during installation, and generate a client secret: Argus links a new installation to your dashboard account only after GitHub confirms you can access it. Repository permissions: Pull requests, Issues and Contents read & write (Contents can be read-only if you don't want @argus-eye fix), Metadata read. Events: pull_request, push, pull_request_review_comment, issue_comment, issues and installation. Generate a private key. The App's slug is the bot's handle: set it as GITHUB_APP_SLUG on the backend and NEXT_PUBLIC_GITHUB_APP_SLUG on the web app (these docs use the default, argus-eye).

  2. Deploy Argus

    Copy backend/.env.example to backend/.env and fill in DATABASE_URL, GITHUB_APP_ID, GITHUB_WEBHOOK_SECRET, the App private key (GITHUB_PRIVATE_KEY_PATH or GITHUB_PRIVATE_KEY), ENCRYPTION_KEY, CLERK_JWKS_URL from your Clerk application, GITHUB_APP_CLIENT_ID and GITHUB_APP_CLIENT_SECRET (the App's client ID and secret; without them the dashboard can't link a new installation), DASHBOARD_BASE_URL (your dashboard's URL, linked from Argus's GitHub comments) and CORS_ALLOW_ORIGIN (your dashboard's origin, so it can call the API). docker compose up from the repo root starts Postgres with pgvector, runs migrations and starts the server; Compose sets DATABASE_URL itself and expects the key at backend/secrets/github-app.pem. Fly.io also works. Run the dashboard from web/ with the Clerk keys and NEXT_PUBLIC_API_URL (your backend) in .env.local.

  3. Install your App on your repos

    Sign in to your dashboard, open Repos and choose Add repos, then install your App on the account that owns it and grant it all repos or specific ones. You can change the selection any time in GitHub. GitHub asks you to authorize the App and sends you back to the dashboard, which checks with GitHub that you can access the installation and links it; the repos you grant appear in the dashboard.

  4. Add your API key and models

    Add your own key on the dashboard's Integrations page (OpenRouter, OpenAI, Anthropic or one of 9 other providers), then pick a model for each pipeline stage in Settings. Argus has no default model, and you pay your provider directly. Memory also needs an embeddings key, or a keyless local endpoint (see Memory storage).

  5. Open a pull request

    With SELF_HOSTED=true on the backend, auto-review is on by default: opened, pushed and reopened PRs are reviewed automatically. With SELF_HOSTED unset or false, auto-review defaults to off: Argus posts a Trigger Argus review checkbox with a cost preview, and ticking it runs the review. A repo or org setting in Settings turns auto-review on or off explicitly under either default. Inline comments carry a GitHub suggestion block when the model supplied a fix.

  6. Teach it your standards

    Choose a review persona, add org rules, or teach patterns with @argus-eye remember. Argus also stores patterns from high-scoring findings, and 👍/👎 reactions and replies from people with write access are stored as feedback for later reviews.

The review pipeline

Every reviewed PR runs through the stages below. Triage, review, scoring and synthesis each take their own model, set per org and overridable per repo.

Review pipeline · stages per run

  1. Triage

    classify + contract

  2. Context

    memory + deps

  3. Review

    one call per file

  4. Score

    dedupe + judge

  5. Synthesize

    suppress + summary

  6. Post & Learn

    comment + memory

Triage, review, scoring and synthesis each use their own model, set per org or per repo, so a cheap model can triage and a stronger one can review. With Deep Review on (off by default), a lead brief runs before review and Pass 2 runs after scoring.
  1. Triage

    Computes the review contract and routes each changed file to skip, skim, security-only or deep using path and size heuristics, which cost nothing. An LLM refines the routing when a code file changes, unless more than 20 files come out deep or security-relevant. An intent-extraction LLM call reads the PR description first. Lockfiles, binaries, and vendored or build output are skipped.

  2. Context

    Adds related files found by GitHub code search, dependents from the code graph (blast radius), stored scenarios for the touched files, SAST hints, and a memory briefing to each file's review prompt.

  3. Review

    One review call per non-skipped file under the 12-rule Review Laws rubric, run in parallel across files. Files triaged security_skim (security-relevant paths, unless the LLM triage pass reclassifies them) get only the security specialist prompt. With Deep Review on (off by default), each deep file gets four specialists instead: bug_hunter, security, architecture and regression. Each attached custom agent adds one more call per deep file, with Deep Review on or off.

  4. Refine

    An LLM judge scores each finding against class-aware thresholds and groups duplicates, whenever a scoring model is configured. Findings within 10 points below their threshold fold into a collapsed Minor notes section; lower ones are dropped. Findings that match something your team dismissed are dropped or downgraded.

  5. Synthesize

    Builds the summary without an LLM: a 1-10 score, one verdict sentence, and a linked row per finding ordered by severity. LLM calls here check the changed files and findings against the PR's stated goal and re-check stored failure scenarios for the touched files against the diff.

  6. Post & Learn

    Posts one GitHub review: the summary plus at most 10 inline comments, blocking first. Findings and learned patterns go to memory, and after posting Argus appends unmentioned changes and grounded Mermaid diagrams to the PR description.

Optionally, a typed classifier sits in front of these stages as a pre-filter. It is not a stage itself: confident answers skip or shrink an LLM call, and everything else flows through unchanged. See TypeSafe Jev.

The review contract

Review depth and posting thresholds depend on the PR's change class.

Before reviewing, Argus computes a contract for the PR: which of 8 change classes it is. Classification is deterministic-first. A hotfix label, then the branch prefix, then path patterns (when more than half the changed files match one class) decide it. When none of those decide, the intent-extraction LLM's class is used at confidence 0.6 or higher; otherwise the PR is treated as production. The class changes file routing and posting thresholds. Draft status, WIP labels and title framing also set depth and evidence-bar labels, which are recorded and displayed but not read by routing. A PR over 60 files or 1,500 changed lines also gets a reduced-confidence note in the summary.

Change classes

production
The default. No routing changes, and the base posting thresholds (critical 35, warning 45, suggestion 55).
migration
Schema and data migrations: migrations/ paths or *.sql files, or cutover/, migrate/, migration/ branches. Every .sql file is reviewed deep, and the critical threshold drops from 35 to 30.
one_time_script
Backfills and one-off jobs: scripts/, tools/, bin/ paths, or spike/, prototype/, poc/ branches. Deep files under scripts/, tools/ or bin/ drop to skim except security-relevant paths; other files keep their depth. Under Deep Review, a script file still triaged deep gets one script reviewer instead of the four specialists. When every reviewed file is under those folders and none is security-relevant, the warning and suggestion bars rise by 10 and 15 and Pass 2 is skipped.
test
Test-only changes: *_test.* files or tests/ paths. Deep test files drop to skim (one pass on the first 100 diff lines) except security-relevant paths; other files keep their depth. Base thresholds apply.
config
Config and infra changes. No path or branch rule assigns it; only the intent LLM does. Routed like production, with the class on record for the judge and the footer.
docs
Documentation: docs/ paths or *.md files. Deep docs files drop to skim except security-relevant paths; other files keep their depth. When every reviewed file is documentation and none is security-relevant, the warning and suggestion bars rise by 10 and 15 and Deep Review's Pass 2 is skipped.
generated
Lockfiles and codegen output (*.pb.go, *_gen.go, dist/). Same routing as docs: deep generated files drop to skim, and when every reviewed file is generated and none is security-relevant, the bars rise and Pass 2 is skipped.
revert
Reverts, detected from revert/ branch prefixes (other revert forms are classified from intent). Routed like production, with the class on record.

Review contract · computed before review

  • draft flag
  • labels
  • branch prefix
  • path globs
  • size

ReviewContract

{ change_class · evidence_bar · depth · signals }

  • file routing

    deep → skim for files whose path matches script/docs/test/generated (not security paths); migration .sql always deep

  • script reviewer

    one reviewer instead of four for script-folder files (Deep Review on)

  • pass-2 eligibility

    scripts/docs/generated skip, if every file matches

  • judge thresholds

    class-aware posting bar, raised only if every file matches

  • glass box footer

    class and depth printed on the review

Deterministic signals set the change class first; the intent LLM fills it only when metadata is silent, and its answer counts only at confidence 0.6 or higher (otherwise production). The class then changes file routing, the script reviewer, Pass 2 and judge thresholds, each only where the changed files' paths back it. depth and evidence_bar are recorded and shown but nothing routes on them. Class and depth are printed in the Glass Box footer.

Depth follows the contract

Routing
For one_time_script, docs, generated and test PRs, a file triaged deep drops to skim only when its own path matches the class and it isn't security-relevant. The class can come from PR text or a branch name, so it never lowers the depth of a file on its own: a production file in a “test” PR stays deep. Migration PRs force every .sql file to deep. Under Deep Review, a one_time_script PR's deep files under scripts/, tools/ or bin/ get one script reviewer instead of four specialists; production code keeps the squad. one_time_script, docs and generated PRs skip Pass 2 only when every reviewed file's path matches the class and none is security-relevant. production, config and revert are routed the same way.
Security paths and migrations
Security-relevant paths keep their triage action under every class. Argus spots them by keywords anywhere in the file path (auth, token, session, secret, crypt, password, permission, rbac, oauth, jwt, cors, csrf, credential, login, signin, signup), not by reading the code, so author.go and tokenizer.ts count too. Migration PRs and PRs touching those paths lower the critical threshold from 35 to 30 for all their findings. Team feedback can still downgrade a security finding one level, but never drop it.
Oversized PRs
Past the soft limit (60 files or 1,500 changed lines by default) Argus reviews at most the 40 highest-risk files without Deep Review and says so in the summary. Past the hard limit (400 files or 20,000 lines) it refuses and posts why, suggesting a split. Both limits are tunable in Settings → Limits.

The contract appears in the Glass Box line of each posted review (inside the collapsed usage block when token usage was recorded) and on the dashboard, e.g. Contract: production/full. See Glass Box & Gauge.

Review laws

A 12-rule rubric appended once to every review prompt. Personas and specialists add focus. The rubric tells the model it overrides them, but a custom persona is free text appended to the prompt and can still conflict with it.

The laws are instructions in the review prompt, not a filter in code, and a model can break them. The judge's prompt does not include them: it scores each finding from 0 to 100 on its own scoring guide, which puts false positives, style-only and linter-catchable findings under 40. Code enforces a few parts of the laws directly: the style cap and the permanent-check exemption from suppression. There is no minimum comment count. A clean PR gets a short summary, not manufactured nitpicks. Six of the laws:

One severity rubric
The same critical/warning/suggestion definitions go into every prompt: every reviewer, persona and change class. Posting thresholds still vary by class (see below).
Silence is a valid review
Zero findings is a complete review. The rubric tells reviewers not to add comments to look thorough.
No praise filler
The rubric forbids praise comments. Praise the model returns anyway can post inline on a changed line. It never becomes a summary row or counts toward the verdict; on a file with a blocking finding it moves to Minor notes.
Style is the linter's job
Style, formatting, naming and import order are never findings under the rubric. Findings in the style or readability category, or worded as style (phrases such as “consider renaming” or “import order”), are capped at a judge score of 30: capped warnings and suggestions are dropped, a critical one falls to Minor notes, and on migration or security-path PRs a critical one can still post.
Evidence before severity
Reviewers must state a concrete failure scenario and file:line before assigning severity; without one, the finding is dropped or at most an FYI suggestion. This is a prompt rule, not a hard gate. The judge's own scoring guide keeps 90–100 for a definite bug, security flaw or regression with clear evidence. A suggested fix is posted only when the model supplied one.
Permanent safety checks
Five checks named in the rubric: destructive SQL with a missing WHERE, secrets or PII entering logs, unit-ambiguous numeric constants, refactors that silently change behavior, unchecked errors. Team feedback can downgrade a matching finding one level but never drop it. Matching is by keywords in the finding text, and class thresholds still apply.

The rest of the twelve: a finding must be a bug a test could assert, a measurable performance or cost problem, a security or data risk, or a break from a documented repo standard (“not how I'd write it” is not a finding); every finding asks for a specific change and supplies the fix; untouched lines are out of scope; a repeated mistake is flagged once, at its root cause; no requests for abstractions the PR does not need today; scripts, docs and generated code get only near-certain, severe findings, while migrations, auth, money and secrets get full depth; and comments are about the code, never the author. The full text is the reviewLaws constant in backend/internal/pipeline/review.go.

Judge scoring

A separate judge call scores each finding 0–100 against thresholds that depend on the change class. It runs on every review that has a scoring model configured. Without one, or when the judge call fails, each finding is scored at its severity's threshold and the score caps below still apply. A missing model shows a setup notice on the dashboard only. After a failed call, the dashboard and the PR comment both say findings were filtered by severity only. Under Deep Review, Pass 2 and cold-file findings get their own judge call.

Class-aware thresholds
Base thresholds are critical 35, warning 45, suggestion 55. One-off script, docs and generated PRs add 10 to warnings and 15 to suggestions, but only when every reviewed file's path matches the class and none is security-relevant; one production file keeps the base bars for the whole review. Migrations and PRs touching security-relevant paths lower the critical threshold to 30. Nothing below the bar is promoted to fill space.
Score caps
After the judge, code caps style findings at 30 (see the style law above) and error_handling findings at 45 unless the file path is security-relevant.
Hard cap: 10 inline comments
At most 10 inline comments per review, blocking first; the rest are counted in a line that links to the dashboard, such as “3 more findings are on the dashboard.” Findings within 10 points below their threshold fold into a collapsed “Minor notes” section in the summary, as do suggestions on a file that has a blocking finding.
Outside the diff
A finding on a line outside the diff has no inline thread. It becomes a summary row that carries its impact, up to 10 such rows; the rest join the dashboard line.
Near-duplicates
When two selected comments are near-identical, the less severe one, or on a tie the lower-scored one, is dropped.

Verdicts lead with “Fix N blocking findings before you merge.”, “Resolve N earlier blocking threads before you merge.”, “Check N warnings before you merge.” or, with only suggestions, “No blocking findings.” A PR with no findings gets a coverage sentence such as “Argus reviewed 8 files in one pass.”, plus a sentence naming any custom agents that ran. Argus posts every review as a comment, never an approval or a change request, so merging stays your call.

Deep review

Off by default. Four specialist prompts review each deep-triaged file in parallel.

With Deep Review on, each file triaged deep gets four specialist calls instead of one pass, all on the review-stage model. When memory is configured, each specialist call can take up to 5 round-trips to use memory tools. A lead-agent call first writes a focus brief per specialist. After scoring, Pass 2 sends files with three or more findings scored 70+ back to the architecture specialist; when it runs, a cold-file security pass then covers up to 5 files with 50+ changed lines and no findings. Findings from those two passes are judge-scored as one batch, then deduplicated. Skim files still get one pass, files triaged security_skim get only the security specialist (the LLM triage pass may move security-relevant files to deep, where they get all four), and skipped files get nothing.

Deep review · opt-in, off by default

Files triaged deep

  • bug_hunter

    logic errors, nil derefs, broken invariants

  • security

    injection, auth bypass, leaked secrets

  • architecture

    error handling, resource leaks, coupling

  • regression

    broken callers, changed behavior, missed edge cases

Ranked findings

max 10 inline, blocking first

With Deep Review on, each file triaged deep gets four specialist calls in parallel instead of one (in a script PR, files under scripts/, tools/ or bin/ get a single script reviewer). Files triaged security-only get just the security specialist, skim files get one pass, and skipped files get none. Specialist findings are merged, deduped and judge-scored; Pass 2 and the cold-file pass run after scoring.
bug_hunter
Logic errors, off-by-ones, nil dereferences, broken invariants, incorrect boolean chains. Told to argue against each bug and report it only if it survives.
security
Injection, auth bypass, SSRF, path traversal, leaked credentials, insecure deserialization. Applies a full attacker model to external entry points and an unexpected-input framing to internal code.
architecture
Error handling design, resource leaks, missing timeouts, coupling and dependency direction, API contract breaks, global mutable state.
regression
For modified code: signature changes that break callers, removed exports or checks, changed response shapes. For new code: missing edge cases, error paths and input validation.

Depth follows the review contract: under a one_time_script contract, deep files under scripts/, tools/ or bin/ drop to skim unless security-relevant, and the script files still triaged deep get one script reviewer (correctness, data safety, idempotency) instead of four specialists; other files keep the four. one_time_script, docs and generated PRs skip Pass 2 when every reviewed file's path matches the class and none is security-relevant. A review over the soft size limit runs without Deep Review. Enable it in Settings → Org defaults → Pipeline features, or per-repo under Repo Overrides. Specialist findings are deduplicated before scoring, and a finding raised by two or more specialists gets +10 on its judge score. Reviewers you write yourself run beside the four as custom agents.

Custom agents

Extra reviewers you define: a name plus a prompt. None are attached by default.

Each attached agent runs one more review call on every deep-triaged file, with Deep Review on or off: beside the four specialists, the script reviewer, or the single pass. Its prompt is appended to the base review prompt as a role overlay, so the Review Laws and the output format still apply. The dashboard labels its findings with its name, and the Glass Box line lists it with the reviewers that ran. Agents don't run on a review reduced past the soft size limit.

Define
On the Custom agents page (/settings/agents, linked from the Custom Agents card on Settings → Repo overrides). A name is up to 80 characters and a prompt up to 8,000; a prompt containing a blocked injection phrase is rejected. An installation holds up to 32 org agents.
Attach
Org-wide through the org default list (“all repos (org default)” on the agents page), or per repo on Repo Overrides. A repo's own list replaces the org list, an empty list detaches every agent, and Inherit org default drops the repo's list.
Repo overrides
A repo agent with an org agent's slug replaces that agent's name and prompt on that repo only, and runs where the org agent is attached. A repo agent with a new slug is repo-only: it always runs on that repo, before the attached org agents.
Run cap
Up to 8 agents run per review: repo-only agents first, then the attached list in order. Agents past the cap stay attached but don't run. Each agent is one more review call per deep file, billed to your key.
Deduplication
With Deep Review off, an agent finding is dropped when a finding already kept on the same file is the same issue and at least as severe; single-pass findings are never removed. With Deep Review on, agent findings go through the same cross-specialist dedup as the four specialists.

A review resolves its agents when it starts, so an edit or a deletion mid-review doesn't change it, and a retry keeps the agents it started with. Deleting an org agent deletes its repo overrides and removes it from every attach list. Agent passes get the specialist memory briefing, not the cross-file context.

Incremental reviews

A push to a reviewed PR re-reviews only the new commits.

When you push new commits to an already-reviewed PR, Argus diffs the last reviewed commit against the new head and reviews only that inter-diff. Earlier findings are not re-posted; they stay in their own threads until auto-resolve confirms a fix or someone resolves them.

Incremental re-review · comment lifecycle on push

Prior comment

git push
  • Flagged code unchangedThread stays open, not re-posted
  • Push fixes the flagged lines (judge-confirmed)Resolved by ‹sha›
  • New issue introducedNew comment posted
When a push triggers a review on a PR that already has a completed review, Argus reviews only the commits since that review and drops new findings that near-duplicate earlier comments. A force-push, a base change or a manual review (@argus-eye review or the trigger checkbox) gets a full review instead. On each push, Argus resolves one of its threads only when a judge model confirms the push fixed it; @argus-eye resolve closes all open threads without that check.
Delta detection
Compares HEAD against the last-reviewed SHA. Only new or modified hunks enter the pipeline. Unchanged files are skipped entirely.
Finding lifecycle
A new finding within 10 lines of an earlier inline finding in the same category is dropped as a repeat. When a push changes flagged lines and a judge confirms the fix (TypeSafe Jev first when enabled, otherwise an LLM on the review-stage model), Argus resolves its own thread with a “Resolved by <sha>” reply. Earlier blocking and warning threads still unresolved on GitHub count toward the score and verdict until someone resolves them.
Cost reduction
Only the inter-diff is sent for review, so token use scales with what changed since the last review rather than with the whole PR.

Incremental mode needs a usable inter-diff from the last completed review. If GitHub's compare of the old and new heads fails or comes back empty, which a force-push or a changed base branch can cause, Argus reviews the whole PR again and says so, and auto-resolve waits for the next push. Force a full re-review any time with @argus-eye review --force.

Auto-resolve limits

Cost
Each re-check the judge model reads is a model call billed to your key. Its spend is recorded on the earlier review whose threads it judges, so it shows on the dashboard, not in the PR's usage table.
Turning it off
Auto-resolve is on by default. Turn it off per org or per repo in Settings, or close every open Argus thread on a PR with @argus-eye resolve (owners, members and collaborators).
Per-push cap
Up to 20 threads are checked per push; the rest stay open until a later push touches them.
Hand resolves
The unresolved count reads GitHub's thread state, so resolving the conversation by hand also clears it.
Timing
If the re-check finishes after the new summary is written, a fixed thread can still be counted until the next push.

What Argus sees

What goes into each review call besides the diff.

Each file's review prompt carries the file's diff (cut to 100 lines when triage marked the file for a skim), the PR title, description and extracted intent, up to 500 lines of the full file, related files, dependents from the code graph, stored scenarios for the file, SAST hints, and a memory briefing. The code graph is parsed from your default branch in the background, not built from reviews.

What the reviewer sees · context fan-in

PR diff

the changed lines

  • Code graphdependents to depth 2, type info · empty until the repo is indexed
  • Related filesup to 5 found by GitHub code search, test names, imports
  • Memoryrepo patterns, past reviews, file synthesis
  • Org rulescustom rules your team wrote in plain language
  • Scenariosknown failure modes attached to these files

Reviewer context

added to each file's review prompt

Each file's review prompt carries the diff, up to 500 lines of the file, and these blocks. Graph blocks stay empty until the repo's code graph is indexed. Specialist calls skip the related-files block, and their memory briefing leaves out org rules and past findings.
Cross-file context
Up to 5 related files (150 lines each), found by GitHub code search on symbol names pulled from the diff with a regex, a guessed test file, and import paths. Primary review pass only. It does not use the code graph.
Blast radius
A background indexer parses the default branch into symbols and call, import and type edges. Files that depend on the changed files, up to depth 2, go into the review prompt, with source for up to 3 direct dependents. Nothing about it is posted on the PR, and name-based resolution misses many cross-package calls.
Scenario memory
Stored failure scenarios for the file (from past critical and warning findings, issues labeled argus or bug, or added by hand) go into the prompt as known issues, and are re-checked against the diff at synthesis.
Static analysis
staticcheck, ESLint and Semgrep results for the PR's main language, as hints the reviewer must check, when the tools are in the image and the PR has 50 files or fewer. The analyzers read the changed files without executing them, and their results are never posted on their own.
Memory briefing
The file's stored summary, up to 5 repo patterns, scenarios and feedback items, and up to 3 org-wide patterns, with dismissals listed as known false positives. Single-pass reviews also get up to 3 org rules and 2 past findings.

Graph-based blocks stay empty until the first index of the default branch finishes. Hotspot notes and bug density need review history. Decision traces are stored for the dashboard and are not read back into reviews.

Architecture visualization

A file-level dependency map of your default branch.

Memory → Files (one repo selected) shows the files in the background index of your default branch, as a sortable list or as a map; no review is needed. On the map, nodes are files, sized by risk and bordered by bug density, and edges are file-to-file calls, inheritance and type use. Reviews add bug density and co-change coupling. Four filters, with a count on each, narrow the list and dim the map.

All files
The default: every file, sorted by a 0–10 risk score, relative to the repo, that blends fan-in, bug density (critical and warning findings per 100 lines), change frequency and strongest co-change coupling.
Choke points
Files with fan-in of 5 or more: five or more other files depend on them.
Hotspots
Files with at least one recorded blocking or warning finding (bug density above zero).
Change together
Files with at least one co-change partner. Coupling is the Jaccard overlap of changed-file lists across the last 200 completed reviews; pairs below 0.3 are dropped.

Navigation

File search
Case-insensitive substring match on file paths. The list shows only the matches, and the map fits them in view. Press / to jump to the search box.
Initial view
The list is sorted by risk, highest first, and any column can be sorted. The map fits all nodes in view on first load. In the list, the arrow keys or j and k move between files and Enter opens one.
File details
Hover a map node for its path, language, risk score and fan-in. Choosing a file, in the list or on the map, opens its details: metrics, symbols, the scenarios watching it, recent findings, the files it changes with, what it depends on and what uses it, matched patterns and its review history.

A push or merge to the default branch queues a re-index, and repos not indexed in 14 days are refreshed. Large repos index in windows of 95 files, so the first index can take a while.

Scenario re-checks

Argus re-checks known failure scenarios against the diff. Nothing is executed.

Scenarios are stored past failures: critical and warning findings from earlier reviews, GitHub issues labeled argus or bug, and ones added by hand. When a PR touches a scenario's files, one LLM call reads the PR title, the scenario and the diff (each file cut to 2,000 characters) and predicts whether the failure still applies. There is no sandbox, no test run, and no new scenario invented from the diff.

Example
earlier findings block

2 earlier findings may still apply to files this PR changes

  • src/billing/cancel.ts · Concurrent cancellations refund the same subscription twice · from acme/web#412
  • src/cache/keys.ts · Recycled user IDs hit the previous holder's cache entry · from acme/web#388

Argus re-checks up to 5 matching scenarios per PR, ranked by how often each came back broken or partial over the last 90 days (never-checked scenarios go first). The review lists only results predicted broken at 0.8+ confidence, collapsed, with the file, the scenario, and the PR that first found it. They never count toward the score or verdict, and nothing shows when no result qualifies. The scenario's details in Memory → Library → Scenarios list its checks with the verdict and its confidence, and the dashboard's review page lists every re-check that review ran. The setting is on when unset and does not require Deep Review: the Simulation and scenarios switch in Settings → Org defaults → Pipeline features.

PR enrichment & Mermaid diagrams

Argus appends unmentioned changes and grounded diagrams to the PR description.

After posting the review, one LLM call compares the PR body with the changed files and findings and lists changes the author did not mention. If the code graph has edges between the changed files, it also asks for up to two Mermaid diagrams. Every node must be a changed file and every edge a parsed graph edge; a diagram that fails that check or the Mermaid parser is dropped. The section sits between marker comments, so a re-run replaces it.

Sequence diagrams
Requested when 3 or more files change. Participants are changed files, and each arrow is a graph edge between them labeled with its edge kind (for example calls), not a traced request flow.
Data flow diagrams
Requested when a finding is in the security category or the changed paths look like auth, token, API or input handling. Same grounding: changed files and parsed edges only, with no type annotations.
Dependency diagrams
Requested when 10 or more files change. Shows graph edges among the changed files only; files outside the PR do not appear.

GitHub renders Mermaid natively. Argus adds at most two diagrams per PR, and none when the default-branch graph has no edges between the changed files. Diagrams go into the PR description, not the review summary. Diagrams need MERMAID_VALIDATOR_BASE_URL on the backend and MERMAID_VALIDATOR_SECRET on both backend and web, or every diagram is dropped. Toggle in Settings → Org defaults → Pipeline features → PR Enrichment.

The conversational review

The summary leads with one verdict sentence, then links each finding to its line.

Every review has three layers: the summary, the inline comments, and the feedback loop.

The summary

Example
review summary

Argus · 7/10 — Hour multiplier is 360,000ms instead of 3,600,000ms

Fix 2 blocking findings before you merge. Argus also found 2 warnings in 8 files.

  • Blocking · Hour multiplier is 360,000ms instead of 3,600,000ms · src/lib/convert/units.ts:L15
  • Blocking · User input passed directly to RegExp without escaping · src/lib/filter/predicate.ts:L42
  • Warning · No NaN check before clamping · src/lib/color/grade.ts:L10
  • Warning · Unbounded bucket array · src/lib/counter/rolling.ts:L28

Dashboard → · React 👎 on an inline comment to dismiss it

Inline comments

Every inline comment follows the same format: severity and category, the issue, why it matters, and a GitHub suggestion block when the model supplied a fix.

Example
inline comment

Blocking · Bug: Double refund on concurrent cancellation

Two concurrent cancellation requests can both pass the status === "active" check. First succeeds at the payment provider, second throws, but the DB update runs for both. No lock or idempotency key guards this check-then-act path.

```suggestion — a GitHub-native suggested change when a fix is clear.

Argus did not compile or run this fix · React 👎 to dismiss

The feedback loop

React 👍 or 👎 on an Argus inline comment, or reply to it. GitHub sends no reaction webhooks, so Argus reads reactions at the next review of that PR or when someone replies. Only reactions from users with write access count, and the majority wins.

Confirm
Records a confirmed outcome, raises the quality score of the pattern the comment matched (if any), and stores a confirmed-feedback memory that later briefings list under Confirmed Findings.
Dismiss
Marks the finding dismissed (the thread stays open) and stores a feedback memory with the PR's change kind. A reason is stored only when a write-access reply explains it. A later finding that matches one dismissal at 0.80+ similarity is dropped; 0.75+ downgrades it one level, and three such matches drop it. A category whose last three outcomes in the repo were dismissed or ignored is dropped too. Security and permanent-check findings are downgraded, never dropped, and dismissals on one-off scripts or prototypes don't apply to production or migration PRs.

Live activity timeline

A running review streams its progress to the dashboard.

While a review is pending or running, the review detail page streams pipeline events over WebSocket, with an 11-step progress bar and an activity timeline.

Live streaming
Shows which file is being reviewed, which specialist is on it, and each finding as it arrives. The client reconnects with backoff and replays missed events.
Scoring results
When the judge finishes, the timeline shows how many findings were kept and dropped, with the cutoffs.
Token & cost counter
Running token and cost totals, updated as triage, review and scoring report usage.
Elapsed timer
Running timer shows total review duration. Auto-scrolls when you're at the bottom, stops auto-scroll when you scroll up to read.

The timeline shows the last 8 rows and expands on request. It renders only while the review is pending or running; after completion the page shows the results instead. A Stop Review button cancels a running review.

Severities

Every finding is tagged with one of four severity levels. Critical, warning and suggestion counts set the 1–10 score, and each has its own posting threshold.

critical
Blocking. Only for evidenced incorrectness, a security risk, data loss or an irreversible change, or a design change that makes the system worse.
warning
Should-fix. Performance problems, error handling gaps, race conditions, or code that works but is fragile.
suggestion
Nit or FYI: meant to be ignorable, and the rubric says not to re-raise it. Style and naming are not findings, and a finding without a concrete failure scenario is at most an FYI suggestion.
praise
The rubric asks for no praise comments. Praise the model returns anyway can post inline on a changed line, but it never becomes a summary row and never counts toward the verdict or score.

Rebalance and the header score

Rebalance
When more than half of a review's findings are blocking, counting nits later moved to Minor notes, the lowest-scored ones are posted as warnings instead. A lone blocking finding therefore posts as a warning.
Header score
Starts at 10 and drops on a log scale with each blocking, warning and suggestion finding, including nits moved to Minor notes and open earlier blocking and warning threads. It never goes below 1, and an unmet stated goal caps it at 7. It is coarse: one blocking finding plus one nit reads 8/10, and a lone finding relabelled as a warning reads 9/10.

Categories

Every finding is also tagged with a category — the type of issue detected.

security
Injection vulnerabilities, leaked credentials, unsafe deserialization, SSRF, path traversal.
bug
Off-by-one errors, nil dereferences, broken invariants, incorrect boolean logic, missing edge cases.
performance
N+1 queries, unnecessary allocations, missing caching, O(n²) where O(n) is possible.
error_handling
Swallowed errors, empty catch blocks, missing error propagation, silent fallbacks.
readability
The fallback for any category the model returns outside the valid set, including style. Readability findings are capped at a judge score of 30. That is below every base threshold, but a critical-labeled one still lands in Minor notes, and on migration or security-path PRs (critical threshold 30) it can post inline.
style
Not a category reviewers may emit: under the Review Laws, formatting, naming and import order are never findings. Style output is re-tagged readability and capped the same way.
type_design
Weak type invariants, stringly-typed APIs, missing generics, poor encapsulation.
testing
Missing edge case tests, brittle assertions, untested error paths, test-only code in production.

Review rules

Rules are org-wide text you write. Each single-pass file review's memory briefing lists up to 10 enabled rules, highest priority first, each cut to 300 characters and about 1,200 characters in all (a rule that does not fit is skipped for shorter ones), so every file gets the same rules. Separately, a finding whose text closely matches a rule is labeled with it; that match needs memory embeddings.

Org rules

Create rules in the dashboard under Memory → Library → Rules. Each rule has a category, free-form content, a priority (low/medium/high), and an enabled flag. Rules are scoped to the installation — they apply to every repo in the org.

Categories: security, performance, style, testing, documentation, error-handling, accessibility, other.

category:    security
priority:    high
content:     Always flag hardcoded API keys or secrets,
             and check raw query strings for SQL injection.

Rules are managed in the dashboard only. Argus does not read config files from your repo.

Model configuration

Each of the 4 LLM stages (triage, review, scoring, synthesis) takes its own provider, model and reasoning effort, from minimal to extra high; a stage calls the endpoint saved with that provider's API key. Triage, review and scoring also take their own max tokens and temperature; synthesis calls use fixed values. Configure org defaults on Settings → Org defaults, override per-repo on Repo Overrides. A repo config wins over the org default, and one review model serves every file in a run. With TypeSafe Jev on, its triage pass can raise a file's review effort above the stage's. There is no built-in default model. Without a review-stage model and key, Argus posts a setup comment and stops. An unconfigured scoring stage skips the judge (findings are filtered by severity only, with a dashboard notice), triage falls back to path heuristics, and synthesis calls fall back to the review model.

What each model stage does, its typical call volume and its tuning lever
StageWhat it doesTypical call volumeTuning lever
triageIntent extraction from the PR description, plus LLM triage when a code file changes and at most 20 files come out deep or security-relevantUp to 2 calls per PRCheap + fast works well
reviewThe findings: one call per non-skipped file; under Deep Review, four specialist calls per deep file (each up to 5 round-trips when it uses memory tools) plus Pass 2 and the cold-file pass; one more call per deep file for each attached custom agent. Also serves pattern and convention extraction, reply analysis, @argus-eye test and the auto-resolve judge (after TypeSafe Jev, when enabled)1 call per file; ×4 per deep file under Deep Review, up to 5 round-trips each; +1 per custom agent per deep file; up to 20 auto-resolve calls per pushSpend your best model here
scoringScores each candidate finding against class-aware thresholds1 call per PRMid-tier is usually enough
synthesisGoal check, PR description, per-file memory, architecture-graph extraction, scenario re-checks, and the Deep Review lead brief and blast-radius check. The summary and verdict are built without an LLMSeveral calls per PR, plus up to 5 scenario re-checksStrong writing quality helps

Supported providers: OpenRouter, OpenAI, Anthropic, Fireworks AI, Groq, Together AI, DeepSeek, Azure OpenAI, AWS Bedrock (limited support), Zhipu AI (GLM), and Vercel AI Gateway. Only providers with a saved API key appear as options. Custom model names are supported — enter any model identifier your provider accepts.

TypeSafe Jev is not on this list and is not a configurable stage. It is a typed classifier that fronts the pipeline, not an LLM provider. See TypeSafe Jev.

TypeSafe Jev

A typed classifier in front of the pipeline: probabilities, not prose. Off by default.

Jev is TypeSafe's “System One” endpoint. Given a state and a set of typed questions (yes/no, pick-from-a-set, or an ordered score), it returns probabilities. It cannot write prose, so it is never an LLM provider and isn't one of the configurable stages above. Its job is cheaper decisions: when Jev is confident (act thresholds of 0.95 to 0.97, by surface), Argus skips or shrinks the LLM call, and the uncertain middle band escalates to your configured models unchanged. It classifies and doesn't compute, so no surface asks it for arithmetic.

Five surfaces consult Jev

Addressed judge
On each push, decides whether an earlier finding was addressed, from the finding text and the file's inter-diff. At 0.97 or higher, with the fix visible in the diff, the thread resolves with no LLM call; at 0.05 or lower it stays open; anything between goes to the LLM judge.
Convention-relation classifier
The memory write gate. Classifies how a candidate convention relates to each similar stored convention (duplicate, refines, contradicts or unrelated) before anything is written. The LLM call is skipped only when every answer is at 0.95 or higher.
Intent verification
Checks the extracted PR-intent block (goal, non-goals, acceptance criteria) against the file list and finding summaries. Jev can only return the all-clear: goal delivered, every criterion met and no finding out of scope, each at 0.95 or higher. Anything else goes to the LLM.
Scoring false-positive pre-filter
Screens every finding against the PR title, author, up to 1,500 characters of PR body and the review-contract summary before the judge spends tokens. A finding rated a false positive at 0.97 or higher, and a concrete defect at 0.5 or lower, skips the judge and is dropped. Reviews with more than 60 findings skip the pre-filter.
Triage shadow
Path, status and about 2,400 characters of raw diff per file, for PRs of up to 40 files (larger PRs skip it). Depth answers are observe-only calibration and never change which files get reviewed. Per-file effort answers can raise a file's review reasoning effort, never lower it; they act at 0.7 confidence, with high picks capped at 1 in 4 reviewed files (minimum one).

Off by default: two ways to turn it on

Server key + opt-in flag
Set TYPESAFE_API_KEY on the backend (TYPESAFE_AI_API_KEY also works) and turn on the jev_classifier feature flag for each installation. Both are required: the flag defaults off and has no dashboard toggle, so the env key alone never sends an installation's data. With no key at either level, Argus makes zero Jev calls.
Or bring your own key
Store a TypeSafe key, and optionally a custom endpoint, in Integrations → TypeSafe Jev. A stored key takes precedence over the env key for your installation and needs no flag: adding the key is the consent to egress.
What is sent
Finding text, PR metadata and a body excerpt, the PR-intent block, diff excerpts for up to 40 files (triage shadow), inter-diff hunks for up to 20 auto-resolve candidates per push, and stored conventions. It goes to api.typesafe.ai, or your stored custom endpoint, for your installation only.

$0.042 per million input tokens; output is free. Answers land in about 200–300ms. The model is pinned to jev-1.13.0, because a moving alias would silently change answers behind tuned thresholds. Setup, the opt-in SQL and endpoint rules: Jev classifier.

API keys (BYOK)

Reviews run on your own provider key. API calls go from the Argus backend to your chosen provider, you pay the provider directly, and the dashboard shows token use per stage. Argus stores review results, PR metadata, memory, and each run's saved state, which includes the diff and the full content of large changed files. It does not clone or store your whole source tree.

Setup

  1. Go to Integrations in the dashboard
  2. Choose a provider (OpenAI, Anthropic, etc.)
  3. Enter your API key. It is stored encrypted.
  4. In Settings, pick a model for each pipeline stage (triage, review, scoring, synthesis)

Security

  • AES-256-GCM at rest — keys are encrypted under your ENCRYPTION_KEY and never stored in plaintext.
  • Decrypted in memory — each LLM call reads and decrypts the stored key again; the embedding provider holding its key stays cached for up to 5 minutes. Keys are not logged.
  • Workspace-isolated — no other workspace can access your keys.
  • Masked — dashboard shows **** plus the last four characters. The full key is never sent to the frontend.

TypeSafe Jev key

The Integrations page also has a TypeSafe Jev card. A key stored there takes precedence over the server's environment key for your installation and can carry a custom base URL. Because adding the key is itself the consent to egress, it doesn't require the jev_classifier opt-in flag. See TypeSafe Jev.

Without a review-stage key and model, Argus posts a setup comment on the PR linking to Settings, once per PR rather than on every push, and does not review.

Memory storage

Where review patterns, conventions and scenario history live.

Memory is stored in Argus's own Postgres database as ordinary rows, each tagged with the embedding space that produced it. Retrieval is a hybrid of vector similarity and full-text search, fused and scoped to your installation. The rows stay in your Postgres, but memory text is still sent to the embeddings endpoint (Voyage by default) and included in prompts to your LLM provider. It stays in-house only if those endpoints are self-hosted too.

What you configure

  1. Go to Integrations in the dashboard and find the Memory embeddings card
  2. Memory embeddings use the provider you configure. The backend defaults to Voyage (EMBEDDINGS_BASE_URL https://api.voyageai.com/v1, EMBEDDINGS_MODEL voyage-4, 1024 dimensions); the shipped backend/.env.example overrides these to the Vercel AI Gateway with voyage/voyage-4-large. Either needs a key for that endpoint: set EMBEDDINGS_API_KEY on the backend or add one on this card. Without a key or endpoint, every similarity-floored memory read returns nothing, so review briefings, rules, rule labeling and dismissal suppression all stop working
  3. Use a Voyage or OpenAI key, or point at any OpenAI-compatible /embeddings endpoint serving 1024-dim vectors; a local endpoint such as Ollama or TEI can run without a key
  4. Embedding settings are per org. Memory itself is kept per repo, plus an org-wide shared space for patterns that apply across repos

Changing the embedding model changes the vector space. A backfill re-embeds existing rows, and retrieval quality can shift until it completes. The similarity thresholds do not re-calibrate on their own: retrieval floors stay at their defaults or your per-install overrides (see Memory tuning), and the suppression floors are fixed.

Review personas

Personas add a prompt overlay that changes the tone and focus of review calls. The Review Laws still apply and take precedence. Set a default per org or per repo; the Settings persona cards list all eight built-in personas and any you have retuned.

default
No overlay: the base review prompt plus the Review Laws.
security_auditor
Security first: injection, auth flaws, secrets in code, input validation, SSRF, crypto misuse. Non-security issues only if critical.
performance_engineer
N+1 queries, missing pagination, hot-path allocations, leaked goroutines or handles, missing caching, O(n²) loops. Non-performance issues only if critical.
mentor
Explains why each issue matters and links to docs or language specs when relevant.
architect
Module boundaries, API contract design, dependency direction, coupling and cohesion.
strict
Asks the model to trace every return path and error branch before calling a file clean. Severity thresholds don't move.
adversarial
Asks the model to assume hostile input and failing network calls on every path. Findings still follow the Review Laws.
fresh_eyes
Reviews as if seeing the codebase for the first time. Flags anything that isn't immediately obvious to a newcomer.
custom
Your own prompt text, appended to the review prompt and cut to 150 characters as the specialist hint. The Review Laws still apply.

The --persona flag (label only)

@argus-eye review --persona strict is accepted, but today it changes only the persona label on the review and progress comment. The review prompt still uses the org or repo default persona.

Auto-review & triggers

Argus supports two trigger modes. Pick per-org, override per-repo.

Auto-review on. Every PR opened, pushed, or reopened is reviewed automatically — no checkbox, no preview. A push to an open PR re-reviews the new commits.

Auto-review off. Argus posts a Trigger Argus review checkbox comment (once per PR) with an estimated token + cost preview. A maintainer ticks the box to run a review on demand.

Default. With SELF_HOSTED=true on the backend, auto-review is on by default. With SELF_HOSTED unset or false, it defaults to off, because reviews cost your tokens. Either way, a repo or org setting turns it on or off explicitly, and SELF_HOSTED never overrides an explicit off.

Precedence
Repo override beats org default. If the repo setting is unset, the org default applies. If both are unset: on with SELF_HOSTED=true, off when SELF_HOSTED is unset or false. Manual triggers (the checkbox and @argus-eye review) run regardless.
Cost preview
The trigger comment shows changed-file count, diff lines, and a historical average of tokens + USD cost across your last 20 reviews for this repo: an average, not a prediction for this PR. USD is omitted when pricing data is unavailable (token-only fallback).
Fallback command
You can trigger a review any time by commenting @argus-eye review, regardless of the auto-run setting; the commenter needs write access. Useful if the checkbox comment is missing (webhook redelivery, PR opened before Argus install).

How the checkbox works

  1. The first PR event (open, push, or reopen) while auto-review is off posts a single comment with a cost preview and - [ ] Trigger Argus review. It posts once per PR; later pushes don't repost.
  2. A user with write access to the repo ticks the box. GitHub fires an issue_comment.edited webhook.
  3. Argus verifies the comment author is your App's bot account (argus-eye[bot] with the default slug) and the ticker has write access (anti-hijack), swaps the checkbox for Running Argus review…, and dispatches the review. A refused click resets the box.
  4. If the pipeline errors, the checkbox is restored with a retry hint. Tick again to run.

Where to toggle

Dashboard → Settings:

  • Org Defaults tab → Auto-review card for the org-wide default.
  • Repo Overrides tab → Auto-review card for a per-repo override.

Gotchas

  • The trigger comment posts once per PR, on the first open, push, or reopen seen while auto-review is off. A PR that predates the install gets it on the next push instead. No comment at all? Use @argus-eye review.
  • Ticking the box on anyone else's comment that mimics our format is ignored — only comments authored by Argus trigger reviews.
  • Only the [ ]→[x] transition triggers a review. Unticking ([x]→[ ]) does nothing, and a running review cannot be cancelled from the checkbox.

Bot commands

Comment on a PR with @argus-eye followed by a command. argus-eye is the default App slug; your install answers to your own App's slug (GITHUB_APP_SLUG), so use that handle in place of @argus-eye below. Argus reacts 👀 when it picks the command up, 🚀 on success and 😕 on failure.

Bot commands · mention → dispatch → effect

@argus-eye <command>

PR comment

dispatch
  • reviewrun a review · --force · --persona
  • rememberstore a pattern · --org for org-wide
  • resolveresolve open Argus threads · owner/member/collaborator
  • fixcommit suggestion blocks to the branch
  • testtest plan · --code drafts test code (not run)
  • helppost the command table
Dispatch parses @argus-eye <command> in a PR conversation comment. review needs write access; resolve needs an owner, member or collaborator; remember needs write access (owner or member for --org). fix and test have no permission check.
@argus-eye review
Trigger a review. Add --force to re-review at the same SHA. --persona is accepted but currently changes only the persona label, not the review prompt. @argus-eye review --force --persona mentor
@argus-eye remember <pattern>
Saves a repo pattern to memory for future reviews. Requires write access; --org saves an org-wide pattern and needs owner or org-member status. @argus-eye remember --org always check for SQL injection in raw queries
@argus-eye resolve
Resolves every open Argus review thread on the PR and marks each finding resolved. Maintainer-only (owner, member, or collaborator). @argus-eye resolve
@argus-eye fix
Commits suggestion blocks from bot review comments (any login ending in [bot], not only Argus) to the PR's head branch as one commit. Overlapping or out-of-range suggestions are skipped. It writes to the branch named like the PR's head branch in the base repo: on a fork PR that fails, or, if the base repo has a branch with the same name (such as main), commits there. It has no permission check: anyone who can comment can run it. @argus-eye fix
@argus-eye test
Posts a test plan built from the latest review's findings: unit, edge case, integration and regression tests. Needs a review on the PR first. Nothing is committed or run. @argus-eye test
@argus-eye test --code
Posts draft test code for the findings, written for the test framework the model detects from imports in the diff. Needs a review on the PR first. Nothing is committed, compiled or run. @argus-eye test --code
@argus-eye review --persona <name>
Accepted, but it currently changes only the persona label on the review and progress comment. The review prompt still uses the org or repo default persona. @argus-eye review --persona strict
@argus-eye help
Lists all available commands and their usage right in the PR. @argus-eye help

Commands run regardless of the auto-review setting — they're explicit intent.

Test generation

Argus can draft a test plan or test code from a review's findings.

One LLM call reads the latest review's findings and the first 5,000 characters of the PR diff, and Argus posts the result as a PR comment. Nothing is committed, compiled or run.

Test plan
@argus-eye test posts a markdown checklist of unit, edge-case, integration and regression tests, based on the latest review's findings on this PR.
Draft test code
@argus-eye test --code posts draft test files for the most critical findings, written for the framework the model detects from imports in the diff. Review the code before adding it.

Test generation does not read memory or the review's context blocks. It uses the org-level review model (repo overrides are ignored) and has no permission check, so anyone who can comment on the PR can spend your key.

Memory & learning

What Argus writes to memory, and what it reads back.

After each review, Argus writes critical and warning findings, suggestions scored 70+, high-scoring patterns, per-file summaries and a PR summary to memory. Praise is kept only as positive feedback, and suppressed findings are not stored. 👍/👎 reactions and replies from people with write access add confirmed or dismissed feedback. Fixes are recorded as merge-time outcomes in Gauge, not as memory. When memory is configured and a review wrote something, the posted review ends with a “Learned” line counting it.

Memory architecture · store · update · retrieve

  1. Each review emits

    findings · patterns · file summaries, then 👍 / 👎 and replies people leave on its comments

  2. Postgres (RAG)

    {repo} + _shared containers · vector + full-text hybrid

  3. Next review retrieves

    file history, patterns, rules, past findings, false positives

Stored as

  • Patterns

    auto-learned code conventions

  • Scenarios

    past findings + argus/bug issues

  • Feedback

    confirms and dismissals from 👍 / 👎 and replies

  • Summaries

    per-file synthesis, PR summaries

  • Confirmed → reinforced
  • Dismissed → dropped or downgraded
  • Files change → scenario flagged outdated

written during each review · read by later reviews

Each review writes its critical and warning findings, suggestions scored 70 or higher, high-scoring patterns and file summaries to memory. 👍 / 👎 reactions (read when the PR is next reviewed or the comment gets a reply) and replies from people with write access record confirm or dismiss feedback. A new finding that matches a dismissal at 0.80 similarity or higher is dropped, or downgraded one level at 0.75; security and permanent-check findings are only downgraded. Each file's review prompt gets a briefing from this memory.
Patterns
Findings scored 80+ (90+ on single-pass reviews) become patterns, an LLM extracts up to 3 reusable patterns per review, and up to 3 conventions are read from each diff's added lines. A convention that contradicts a stored one opens a conflict for a maintainer to settle. You can add or delete patterns with @argus-eye remember, the dashboard or MCP.
Scenarios
Stored failure cases from three sources: critical and warning findings from reviews, GitHub issues labeled argus or bug, and ones added in Memory → Library → Scenarios. Each holds a description, a severity and the files it applies to. A review marks scenarios on its changed files as outdated, except the ones it just re-checked. Deactivate one from its details there; reviews stop checking it.
Decision traces
One Postgres row per review finding that suppression doesn't drop: file, severity, text and PR. The dashboard counts the PRs behind them per file (Memory → Overview and Memory → Files), so a re-reviewed PR counts once. A file's details list its newest ones. Traces are not read back into reviews.
Context graph
The code graph parsed from your default branch: files, symbols, and call, import and type edges. Memory → Files joins it with each file's patterns, recent findings and decision traces.

How feedback changes later reviews

A 👍 or 👎 on a comment that matched a pattern moves that pattern's quality score, and a 👎 can drop or downgrade similar findings later. When a pattern's quality falls below 0.4, the judge is told to score matching findings lower. Scenarios with no check in the last 90 days are re-checked first, then those that most often came back broken or partial. Org-wide patterns stop reaching briefings after about 128 days without being re-learned.

Memory rules

Each signal that changes memory, who counts for it, what Argus does with it, and where it stops.

👍 or 👎 on an Argus inline comment
Who counts: People with write access. Bot accounts are skipped.What Argus does: 👎 records a dismissal of that finding and 👍 a confirmation. Removing the reaction retracts its memory entry.Where it stops: GitHub sends no reaction events, so Argus reads reactions at the next review of that PR, or when someone replies to that comment. A 👎 carries no reason.
A reply in an Argus thread
Who counts: Anyone who replies gets an answer. Only replies from people with write access change memory.What Argus does: The review model answers in the thread and picks resolve, clarify, stand firm, or step back for this kind of change. A resolve with a lesson specific to the repo stores a dismissal, with the reply as the developer explanation, and stores the lesson as a pattern.Where it stops: The model decides whether the reply was a dismissal. There is no confirm step.
A new finding close to a stored dismissal
Who counts: Automatic.What Argus does: Dropped before posting at similarity 0.80 or higher to one dismissal, or when three dismissals match at 0.75 or higher. From 0.75, it posts one severity lower with “— Previously dismissed a similar finding (PR #N).”Where it stops: Security findings and the five permanent checks (destructive SQL, secrets or PII in logs, unit-ambiguous constants, behavior-changing refactors, swallowed errors) are lowered at most one level, never dropped; the checks are keyword matches on the finding text. Dismissals on one-off scripts or prototypes don't apply to production or migration PRs. The PR shows only the count of dropped findings; the dashboard's review page lists each one with its reason.
The last three outcomes in a category were dismissed or ignored
Who counts: Automatic. A finding nobody acted on before merge counts as ignored, and so does one where the model answered a writer's reply with clarify.What Argus does: New findings in that category are dropped for the repo.Where it stops: Security and permanent-check findings still post. No screen lists these categories yet; a newer outcome that is not negative ends the streak.
Conventions in the diff's added lines
Who counts: Automatic.What Argus does: Up to three per review, read from at most 2,000 characters of added lines and compared with stored conventions as a duplicate, a refinement or a contradiction.Where it stops: A convention is a model's reading of added lines. A contradiction posts a Convention conflict comment, and until someone with write access ticks one side, reviewers' memory briefings list both conventions under “Disputed — do not enforce”. The losing convention stops being retrieved.
A finding the judge scores highly
Who counts: Automatic.What Argus does: Stored as a pattern at a score of 90 or higher in a single-pass review, 80 under Deep Review; without a scoring model, or when the judge call fails, blocking and warning findings are stored instead. A later finding above 0.80 similarity cites its PR: “— Matches a prior fix in PR #N.”Where it stops: A high judge score is the model's rating, not a person's confirmation. “A prior fix” is the label for that pattern; Argus does not check that a fix merged.
@argus-eye remember <pattern>
Who counts: Write access. --org needs an owner or organization member.What Argus does: Stores a pattern for the repo, or org-wide, and replies “Remembered (this repo): …”.Where it stops: 30 days after an org-wide pattern was last written, it starts losing 0.05 confidence a week, and it drops out of briefings below 0.30.
retire_memory, delete_memory, the dashboard's Memory → Library → Patterns
Who counts: MCP tokens with argus:memory:write, and dashboard users.What Argus does: Stop a memory being retrieved, or delete a pattern.Where it stops: Retiring cannot be undone and needs a reason. Memories the pipeline learned need an explicit confirm flag.

Every similarity rule above needs an embeddings endpoint (see Memory storage). Without one, rows are still written, but similarity search returns nothing, so dismissals stop dropping findings and citations stop appearing. The suppression and citation floors are fixed. Settings → Memory adjusts three retrieval floors (a fourth, Failure recognition, is unused) and can switch off org-wide decay (see Memory tuning).

Glass Box & Gauge

Each posted review states how it ran.

The Glass Box line on each posted review states the contract, which reviewers it was set to run, how many findings team feedback suppressed, and how long the review took. When token usage was recorded, it sits inside the collapsed usage block above a per-stage table of model, tokens and cost; otherwise it stands alone. Every token it lists is billed to your key.

Example
glass box line

Contract: production/full · checked: bug_hunter, security, architecture, regression · 2 suppressed by team feedback · review took 1m42s

Contract
The change class, then the recorded depth. The class sets routing and the judge's bar.
checked
Follows settings, not what ran. single-pass review means one review call per file triage did not skip; with Deep Review off, a security-relevant file still gets the security prompt, and this still reads single-pass review. With Deep Review on it lists all four specialists (script_review on one_time_script PRs) even if no file was triaged deep. Custom agents follow, by name, only when they ran: a file was triaged deep and the review was not reduced.
suppressed
Findings dropped by dismissal memory only. Findings the judge scored under the bar, or dedup merged, are not counted on the PR.
took
Time from the start of the run until the summary was written.
Usage table
One row per stage that ran before the summary was written, with its model and tokens; the cost column is left out when no cost is known. Later spend (memory writes, PR description edits, re-checks on push) shows on the dashboard, not in this table.

Gauge — address-rate telemetry

Gauge records, when a PR closes, whether the code near each posted finding changed before merge. Results are on the Stats page; nothing is posted to GitHub.

Gauge · did the comments change the code

PR closes

merged or not

diff reviewed commit → final head · ±3 lines
  • addressed_humanfull weight

    a human commit touched the flagged lines

  • addressed_agent×0.5

    a bot-pattern author touched the flagged lines

  • ignoredzero

    merged with the flagged lines untouched

  • deferredzero

    PR closed without merging

vw_review_gauge

address rate per category × change class

When a PR merges, the gauge diffs each reviewed commit against the final head and records one outcome per posted finding (up to 100): addressed if code within ±3 lines of it changed, otherwise ignored. A PR closed without merging marks every finding deferred. It is a proximity heuristic, not proof of a fix. Outcomes feed the dashboard's Review Gauge, and ignored outcomes count toward category auto-suppression.
Address rate
A finding counts as addressed if commits pushed after the review changed code within 3 lines of it. Untouched findings on merged PRs count as ignored, and findings on unmerged PRs as deferred. Human fixes weigh 1.0 and agent fixes 0.5, where agent means the last commit to that file came from a login ending in [bot], -agent, -bot or _bot (or listed in ARGUS_AGENT_LOGINS). It is a proximity heuristic, not proof the bug was fixed.
Per category, per change class
Address rate is broken down by finding category and change class, so you can see which kinds of findings your team acts on.

Insights & risk

Which files keep drawing findings.

With one repo selected, Memory → Overview lists the files with the most findings, and Memory → Files adds a findings column to every file. The Stats page covers scores, cost and Gauge across repos.

Hot files
Files ranked by how many PRs recorded review findings on them in the last 90 days. A re-reviewed PR counts once.
Risk scores
Memory → Files gives each file a 0–10 risk score, relative to the repo, that blends fan-in, bug density, change frequency and coupling. It doesn't change how a PR is reviewed; separately, files with fan-in of 5+ or 3+ past bugs get a choke-point or hotspot note in their review prompt.
Findings per file
A file's details in Memory → Files list the newest findings recorded on it, with their severity and a link to the review.
Quality trends
The Stats page charts the average review score per day and breaks it down by repo and by PR author.

Token & cost tracking

Tokens and cost per stage, for each review.

Argus records per-stage token usage and cost for every review. Model and provider are tracked independently for each stage. Token data persists on failed and cancelled reviews too. Cost is the provider-reported value when there is one, otherwise an estimate from price tables, and $0 for a model with no known price.

Per-stage breakdown
These stages record input tokens, output tokens, model, provider, and cost: intent, triage, review (per-file, per-specialist and per custom agent), lead agent, graph, file synthesis, scoring, enrichment, conventions, patterns, acceptance, cross-PR, scenario re-checks, and auto-resolve. Auto-resolve judge spend is recorded on the earlier review whose threads it judges, whether or not it closes them. TypeSafe Jev calls are recorded in the stage they front-run (triage, scoring, intent, conventions, auto-resolve) under provider typesafe. Reply analysis and @argus-eye test calls are not attached to a review.
TokenPill
Hover any TokenPill in the review detail page to see the full cost breakdown per stage, including model name and provider.

Light mode

The dashboard has a light theme.

Toggle between dark and light with the Sun/Moon icon in the sidebar footer or Cmd/Ctrl+Shift+D. Your choice is saved in localStorage and applies without a page reload.

Toggle
Click the Sun/Moon icon in the sidebar footer (shown when the sidebar is expanded) or press Cmd/Ctrl+Shift+D. No page refresh required.
Default theme
The dashboard starts dark until you toggle it. It does not read your OS prefers-color-scheme setting.
Coverage
Dashboard pages, including the architecture graph and code diffs, support both themes; the graph uses a warm cream palette in light mode. Marketing and sign-in pages stay dark.

Feature flags

Toggle capabilities per-org from the dashboard.

These flags are scoped to the whole installation and take effect on the next review. Find them under Settings → Org defaults → Verification features. Per-repo pipeline toggles live under Settings & controls.

Issue Acceptance
on by default Checks each linked issue’s acceptance criteria against the diff with one LLM call per issue (up to 5) and adds an Issue Coverage section, with a verdict per criterion, to the review summary on the dashboard. Links come from GitHub’s closing keywords and Development panel, plus close/fix/resolve/refs mentions in the PR body. Informational only.
Cross-Repo PR Checks
on by default For PRs linked in the description, one LLM call looks for risks that appear only when the changes combine (schema races, contract drift, deploy ordering and 6 other categories) and writes them into a section of Argus’s review. When 2+ linked PRs share an issue, it also judges joint issue coverage. Informational only.
Max Linked PRs
default 5 Caps how many linked PRs the cross-PR worker fetches per review. Any integer 1–20.

Existing installations keep whatever was stored before the defaults flipped; the on-by-defaults apply only where nothing was saved. The jev_classifier flag has no toggle here; see TypeSafe Jev. Deep dives: issue acceptance, cross-repo PR checks, memory tuning, FAQ.

Settings & controls

These pipeline toggles are set per org and can be overridden per repo.

Auto-review
off by default · on with SELF_HOSTED=true Review every PR automatically when it is opened, pushed or reopened; a push re-reviews the new commits. When off, Argus posts a Trigger checkbox (once per PR) with a token/cost preview, and a user with write access ticks it to run a review.
Auto-resolve
on by default On every push, finds Argus threads whose flagged lines changed (within ±3 lines) and asks a judge, up to 20 candidates per push, whether each was actually fixed. With TypeSafe Jev enabled, Jev answers first and settles confident cases without an LLM call; the rest go to an LLM judge on the review-stage model. Confirmed fixes are resolved with a “Resolved by <sha>” reply. Runs even when auto-review is off, and LLM judge calls bill to your key.
Deep Review
off by default Four specialists (bug_hunter, security, architecture, regression) on each deep-triaged file, plus a lead brief, Pass 2 on files with several high-scoring findings, and a cold-file security pass.
Custom agents
none attached by default Attached agents each add one review call per deep-triaged file, with Deep Review on or off. The org default list applies unless a repo sets its own. See Custom agents.
Cross-File Context
on by default Adds up to 5 related files (150 lines each) to the primary review call, found by GitHub code search on symbols from the diff, a guessed test file, and import paths. It does not use the code graph.
Blast Radius Analysis
on by default Adds files that depend on the changed files (up to depth 2, from the default-branch code graph) to each review prompt, with source for up to 3 direct dependents. Nothing about it is posted on the PR.
Simulation and scenarios
on when unset Re-checks up to 5 stored failure scenarios for the touched files against the diff, one LLM call each; nothing is executed. The same switch controls scenario memory: storing scenarios from findings and adding known issues to review prompts. It works with or without Deep Review.
PR Enrichment
on by default Appends changes the description doesn't mention, plus up to two graph-grounded Mermaid diagrams, to the PR description.
Pattern Learning
on by default An LLM extracts up to 3 reusable patterns per review from findings scored 75+ (80+ on single-pass reviews).
Convention Learning
on by default An LLM extracts up to 3 conventions from each diff's added lines and checks them against stored ones for duplicates and contradictions.
File Synthesis
on by default An LLM writes a summary of up to 200 words for up to 10 files with strong findings; later reviews of those files read it back.
Architecture Graph
on by default Runs an extra LLM call per review that extracts components from the diff and stores them. Nothing reads that output today; blast radius uses the parser-built code graph, which is indexed independently of this toggle.

Pipeline toggles are per-org on Settings → Org defaults and overridable per-repo on Repo Overrides. The Limits tab caps review size. Past a soft limit Argus reviews at most the 40 highest-risk files without Deep Review; past a hard limit it refuses and posts why (defaults: 60/400 files, 1.5k/20k changed lines, 3M/10M average tokens; the token measure only engages once the repo has review history). Changes take effect on the next review.

MCP server

Use Argus's memory and reviews from your own agent.

Argus exposes team memory and review results over the Model Context Protocol. Any MCP client that supports remote servers with OAuth (Claude Code, Cursor, Claude Desktop) can connect to your Argus API origin plus /mcp, the URL you set as MCP_RESOURCE_URL. One browser login, then pick one organization: the connection sees only that org's repos and memory. Switching orgs means authorizing again. There are no API keys: the server accepts only OAuth access tokens, and refreshing them is up to your MCP client.

MCP server · your agent ↔ Argus memory and reviews

  1. MCP client

    Claude Code · Cursor

  2. discover

    GET /.well-known/oauth-protected-resource

  3. authorize

    Clerk OAuth · pick ONE organization

  4. call

    POST /mcp · bearer verified (iss · aud · scopes)

argus:read

  • list_repos
  • search_memory
  • get_memory_briefing
  • list_reviews
  • get_review_status
  • get_review

argus:memory:write

  • create_memory
  • delete_memory
  • retire_memory
  • start with list_repos → repo_id / installation_id
  • retire_memory stops influence · delete_memory removes one record
  • confirm_pipeline_learned guards learned memory
An MCP client that supports remote servers with OAuth connects to your Argus API origin plus /mcp, the URL set as MCP_RESOURCE_URL. The route is off by default: enable it with MCP_ENABLED=true plus CLERK_JWKS_URL, CLERK_ISSUER_URL and an https MCP_RESOURCE_URL ending in /mcp; the backend will not start if one is missing. One browser login picks one organization per connection, and the token sees only that org's repos, memory and reviews.
Scoped by grant
argus:read covers every read tool; argus:memory:write covers create_memory, delete_memory, and retire_memory. A read-only grant cannot mutate memory, whatever flags a call sets.
Guard rails on writes
Deleting or retiring a memory the review pipeline learned requires confirm_pipeline_learned=true, and neither can be undone. Writing an org-wide memory requires confirm_shared=true; that memory counts as human-authored, so it can later be retired or deleted without a confirm flag.
retire beats delete
To stop a memory influencing reviews, use retire_memory. delete_memory removes one contributing record, and the memory may remain searchable.

Start every session with list_repos — repo IDs feed every other tool. Full tool reference: /docs/mcp. The route is off by default and 404s while disabled. To turn it on, set MCP_ENABLED=true, CLERK_ISSUER_URL, CLERK_JWKS_URL, and MCP_RESOURCE_URL (an https URL ending in /mcp) on the backend. Setup guide: README § MCP.