Skip to content

Changelog

Release notes for Argus.

  1. The front-runner

    • TypeSafe Jev: a typed classifier now fronts five decision surfaces — the addressed judge on push, the convention-relation memory write gate, intent verification, the scoring false-positive pre-filter, and a triage shadow. It answers typed questions (yes/no, choice, score) with probabilities in ~200–300ms; it can't emit prose and never acts as an LLM provider.
    • Confident answers act, uncertain ones escalate: at 0.95–0.97+ confidence (by surface) Jev's answer skips or shrinks the LLM call; the middle band flows to your configured models unchanged. Jev pre-filters stages — it isn't one.
    • Off by default, two gates to egress: the server-key path needs both a TYPESAFE_API_KEY on the backend and a per-installation jev_classifier opt-in — the env key alone never sends tenant data.
    • Per-installation BYOK: store your own TypeSafe key (and optional base URL) on the Integrations page. A stored key takes precedence over the env key and needs no flag — adding the key is the consent to egress.
    • The triage shadow classifies path, status, and ~2,400 chars of raw diff for PRs of up to 40 files (larger PRs skip it). Its depth answers are observe-only calibration; its per-file effort answers can raise a file's review reasoning effort (never lower it), with high picks capped at 1 in 4 reviewed files (minimum one); effort picks act at 0.7 confidence.
    • $0.042 per million input tokens, output free. Model pinned to jev-1.13.0 so answers don't drift behind tuned thresholds.
  2. The review doctrine

    • Review Contract: every PR gets a computed contract with one of 8 change classes (production, migration, one-off script, test, config, docs, generated, revert), picked from a hotfix label, the branch prefix, or the changed paths. The intent LLM picks the class only when those are silent. Visible on every review.
    • Routing follows the change class. Docs, test, generated and one-off-script PRs have their deep files downgraded to skim (security-relevant paths keep their depth), and migration PRs get every .sql file reviewed deep. With Deep Review on (off by default), the deep files left in a one-off-script PR get a single balanced reviewer (correctness + data safety) instead of the four specialists, and docs, generated and one-off-script PRs skip the second pass.
    • Review Laws: one severity rubric in every review prompt, silence is a valid review, no praise comments, style is the linter's job. The rubric asks every finding for a concrete failure scenario, file:line evidence, and a suggested fix.
    • Judge scoring now runs on every review that has a scoring model configured, with class-aware thresholds; near-threshold findings fold into a collapsed Minor notes section. Posting caps inline comments at 10. No minimum-comment behavior.
    • Team-feedback memory: dismissals become semantic memories with the change kind (a reason is stored only when a reply from someone with write access gives one). A single close match to a dismissal drops a finding, three or more weaker matches also drop it, a single weaker match downgrades it one level, and a category is auto-suppressed after 3 negative outcomes in a row. Team feedback can downgrade a security finding but never drops it. One-off-script dismissals don't silence production reviews.
    • Re-reviews resolve Argus's own fixed comments (“Resolved by <sha>”) and only post what's new.
    • Glass Box footer on every review: contract, what was checked, suppressed count, duration. Gauge tracks address rate per category per change class: whether code within 3 lines of each posted finding changed before merge, with agent fixes counted at half weight.
    • Reviews of PRs over 1,500 changed lines or 60 files carry a reduced-confidence note and a split recommendation.

Get new entries by email

One email when a changelog entry or a blog post goes up.

The form posts to Buttondown, which sends the emails and handles unsubscribes.