Skip to content

Memory tuning

Last updated

Argus remembers things across reviews. Confirmed patterns, known scenarios, dismissed findings, per-file synthesis — all of it feeds back into future reviews as context. The match is semantic, not exact text: a similarity score between 0 and 1 gates whether each memory influences a review.

Three of those gates are tunable per-org, plus one switch for how org-wide patterns decay. A fourth setting, scenario_trigger, is no longer used. Leave the defaults alone unless you see a specific failure mode in your reviews.

specialist_min filters the reviewer briefing on every review, and the Deep Review specialists' briefings when Deep Review is on (it is off by default).

Where to tune

Open Settings from the dashboard sidebar and switch to the Memory tab. Changes apply to the next review. An overridden field shows an amber border and a delta chip; reset to default via the circular-arrow icon next to each control.

Thresholds

Each gate is a similarity cutoff in [0, 1]. Higher = stricter (fewer but more relevant matches). Lower = more permissive (more context, more noise). Scores are raw cosine similarity over the configured embedding space — they sit high, so the defaults are higher than intuition suggests. An explicit 0 turns that floor off, except for scenario_dedupe, where 0 falls back to the default.

finding_enrich

Default 0.70. The floor for per-finding memory lookups (the pattern and rule matched to each finding, and the dismissed-finding lookup) and for the past-review section of the reviewer briefing. The most recall-oriented gate: a miss costs context, a false hit costs one noisy line. A rule match at this floor adds "Matches a repo rule." to the posted comment. Pattern citations on a comment, such as "Matches a prior fix in PR #N", need a match above a fixed 0.80, and repeats of dismissed findings are dropped or downgraded by fixed floors (0.80 / 0.75), so lowering this gate adds neither. Raising it above 0.75 has a cost: the dismissed-finding lookup stops returning matches between 0.75 and the new floor, so those repeats are no longer downgraded or counted toward the three-match drop. At 0.80 or higher, only security findings and permanent checks still get downgraded.

  • Raise (e.g. 0.80) if comments cite repo rules, or briefings pull in past reviews, unrelated to the change: the match is too loose.
  • Lower (e.g. 0.60) if you have a mature pattern library but briefings rarely include any of it: the gate is too strict.

specialist_min

Default 0.80. Server-side similarity cutoff for the pattern sections of the reviewer briefing: the repo's patterns, scenarios and feedback for each file, and org-wide patterns. The default single-pass reviewer gets this briefing, and so do the Deep Review specialists (bug hunter, security, architecture, regression) when Deep Review is on, and every custom agent pass. Stricter than finding_enrich because irrelevant matches use up prompt budget.

  • Raise (e.g. 0.90) if briefings feel noisy, with irrelevant past findings diluting the signal.
  • Lower (e.g. 0.70) for small repos where the pattern library is still thin and you want the briefing to reach further.

scenario_trigger

Unused. Argus no longer reads this setting (the Failure recognition slider); a saved value is kept but has no effect. The "Triggered N times" count on the dashboard's scenarios list goes up once each time a scenario is re-checked, whatever the verdict, and does not change which scenarios get re-checked.

scenario_dedupe

Default 0.95. When a new candidate scenario is extracted from a review's findings, an existing scenario at or above this similarity counts as a duplicate and the new one is skipped. It sits in the top decile on purpose — genuine duplicates score near 1.0, and a lower bar silently merges distinct scenarios.

  • Raise (e.g. 0.98) if you're seeing distinct scenarios silently merged.
  • Lower (e.g. 0.90) if your scenarios list has obvious duplicates accumulating.

Org-wide pattern decay

Some patterns apply to every repo in an installation and live in a shared container: auto-learned patterns that don't reference repo-specific file paths, patterns added with @argus-eye remember --org (use your App's slug, GITHUB_APP_SLUG, in place of argus-eye), and org-wide patterns written from the dashboard or MCP. Learnings from replies to Argus comments stay in the repo they came from. Without decay, one bad shared pattern would stay eligible for every repo's reviewer briefing indefinitely.

Argus computes a shared pattern's confidence from its age at read time, each time a briefing is assembled. Nothing is written back and nothing is deleted:

  • Day 0–30: grace window, full confidence 1.00, no decay.
  • Day 30+: confidence drops 0.05 per week since the pattern was last written, always measured from 1.00 (never compounded).
  • Below 0.30 (~14 weeks past grace, ≈4 months) the pattern drops out of the org-wide section of the reviewer briefing. It stays stored. Per-finding citation lookups and MCP search_memory don't apply this floor, so it can still match there.
  • Re-learning (the pipeline extracts the same pattern again) rewrites the row, which resets confidence to 1.00 and restarts the clock. Adding an identical org pattern again by hand does not, because the memory mirror leaves an existing live row as it is.

disable_shared_decay: the Confidence decay switch on the Memory tab. Switching decay off stores disable_shared_decay=true, which keeps every shared pattern at full confidence until someone retires or deletes it. Default: decay on. Useful when you want a person, not age, to take org-wide patterns out of reviews, or early in a rollout while you watch which patterns accumulate.

How to know it's working

finding_enrich and specialist_min are applied server-side as similarity cutoffs on the memory search itself, so they don't emit a per-check log line. Their effect is mostly invisible in review output: both shape the reviewer briefing, which is not posted, and pattern citations on a comment use a fixed 0.80. The most direct visible signal is the "Matches a repo rule." tag, which follows finding_enrich.

Safe defaults

If you're unsure, don't tune. The defaults were set against similarity distributions measured on a live review corpus embedded with voyage-4-large. A different model shifts the distribution. Start with defaults, watch a few weeks of reviews, then tune only the specific gate that's misfiring.

Memory embeddings come from the provider you configure. If EMBEDDINGS_BASE_URL and EMBEDDINGS_MODEL are unset, the backend uses voyage-4 (1,024 dimensions) from https://api.voyageai.com/v1, which needs a Voyage key. backend/.env.example instead ships the Vercel AI Gateway at https://ai-gateway.vercel.sh/v1 with voyage/voyage-4-large, which needs a gateway key. Either way, set EMBEDDINGS_API_KEY or add an embeddings key on the dashboard's Integrations page (Memory & classifiers → Memory embeddings).

The dashboard sets these per-installation — all repos under one installation share the same gates. The repo-level settings schema accepts the same keys if you need a per-repo override via the API.