What is Code Simulation?
Last updated
Code simulation means working out what a code change does under specific conditions, such as increased load, network failures, race conditions, or edge-case inputs. Some tools execute the change in a sandbox; others only reason about it without running anything, which is a prediction rather than a test. It's the review equivalent of asking 'what breaks if this changes?'
Why does code simulation matter for engineering teams?
Many bugs that reach production are logic failures that show up only under specific conditions, which a happy-path read of the diff can miss. Running the change under those conditions observes the failure directly. Reasoning about it without execution can point a reviewer at the risk, but the prediction can be wrong.
How does Argus handle code simulation?
Argus does not execute code. What Argus calls simulation is a re-check of stored failure scenarios: critical and warning findings from earlier reviews, GitHub issues labeled argus or bug, and scenarios added through the API. When a PR changes files a scenario covers, Argus picks up to 5 and makes one LLM call for each, with the PR title, the scenario text, and each changed file's diff truncated to 2,000 characters. The model predicts whether the scenario still breaks. Only confident 'broken' verdicts appear on the PR, as earlier findings that may still apply. Argus does not create scenarios from the new diff, so files with no stored scenario get no check. Tools that run the PR in a sandbox, such as Greptile's TREX, observe runtime behavior; Argus does not.