Run an evaluation for an agent (preview)

[This article is prerelease documentation and is subject to change.]

After you create a test set with conversations, run an evaluation to measure your agent's performance. The evaluation processes each conversation and produces scored results based on the selected test method.

Important

Prerequisites

To run an evaluation for an agent:

  1. Open your agent in Copilot Studio.
  2. Select the Evaluate tab.
  3. In the Evaluation dropdown, select the evaluation you want to run.
  4. Add at least one conversation to the test set if you haven't already.
  5. In the Configure test set panel, verify:
    • The evaluation Name is set.
    • The Test method is configured (for example, General quality).
  6. Select Evaluate to start the evaluation. The evaluation processes each conversation in the test set. Depending on the number of conversations, this process might take several minutes.

Tip

Run the same evaluation multiple times. Each run is saved separately, so you can compare results across runs and see how changes to your agent affect quality.

Re-run an evaluation

After you change your agent's instructions, knowledge, or tools, re-run the evaluation to measure the impact:

  1. On the Evaluate tab, select the evaluation you want to re-run from the Evaluation dropdown.
  2. Under Recent results, select Evaluate test set again or select the Evaluate test set icon within the test set to start a new run.
  3. Compare the new run's results with previous runs to see if the changes improved or degraded performance.