Question 275
Your team is designing an agent evaluation framework. They need to test whether the agent handles tool failures gracefully across 50 different error scenarios. Running each scenario manually takes 10 minutes. What is the most efficient approach?