Microsoft has released run-assert-eval, an open-source tool designed to measure where AI agents fail and how often those failures occur.
It identifies relevant risks, evaluates agent behavior against them and can generate runtime policies based on the results. Microsoft says the tool can then rerun the same evaluations to verify whether the applied fixes actually reduced the failures.