What is Failure Scenario Testing?
Last updated
Failure scenario testing is the practice of analyzing code changes by reasoning about what happens under adverse conditions (network timeouts, concurrent access, data corruption, resource exhaustion) rather than just the happy path. It answers the question: 'What breaks if this changes under load?'
Why does failure scenario testing matter for engineering teams?
Production failures rarely happen under normal conditions. They happen when a cache expires during a traffic spike, when a downstream service returns unexpected data, or when two threads hit the same race condition. Reviewing only the happy path means missing the scenarios that actually cause outages.
How does Argus handle failure scenario testing?
Argus does not generate or run failure scenarios from the diff. It keeps stored scenarios: critical and warning findings from earlier reviews, GitHub issues labeled argus or bug, and scenarios added through the API. When a PR changes a file a scenario covers, an LLM re-checks up to 5 of them against the diff text and predicts whether each still breaks; nothing is executed. Confident 'broken' verdicts appear in the review as earlier findings that may still apply. Separately, the Review Laws rubric asks every finding to state a concrete failure scenario and file:line before a severity is assigned. That is a prompt rule backed by the judge's scoring, not a hard filter.