---
sourceDocument: Australia Enable AI
sourceDocumentLink: https://servicenow-prod.fluidtopics.net/r/pt-BR/intelligent-experiences

 Release :

    - australia

ft:locale :

    - pt-BR

ft:publication_title :

    - Australia Enable AI

ft:clusterId :

    - platai

bundleId :

    - platai

workflow :

    - Platform


---

# Evaluate agentic AI

# Evaluate agentic workflows and AI agents {#ariaid-title1}

* Versão de lançamento: Australia
* 
* Atualizado 12 de mar. de 2026
* 
* ![](https://www.servicenow.com/docs/portal-asset/ico-clock) 4 min. de leitura

Evaluate agentic workflow and AI agents against datasets of your choice to monitor performance and evaluate against different benchmarks.

## Antes de Iniciar

Evaluation runs require execution log data of the agentic workflow or AI agent you want to evaluate. You can create execution log data by testing in AI Agent Studio, triggering agentic AI in Now Assist in Virtual Agent or Now Assist panel. You can also create execution log data after setting up your evaluation run.

For more information about testing agentic workflows, see [Manually test the execution of an agentic workflow](https://servicenow-prod.fluidtopics.net/EyksAdiTjKwRc~m2JvT7uQ "Test your agentic workflow in AI Agent Studio to analyze how it functions while it executes the instructions that you defined.").

For more information about getting started with agentic evaluations, see [General guidelines for agentic evaluation runs](https://servicenow-prod.fluidtopics.net/usbMSThBd5JXdnav74g3Yw "Learn about agentic evaluation runs and different recommendations for evaluating your AI agents and agentic workflows against datasets to check for completion, performance, and tool execution.").

Role required: sn_aia.admin

## Procedimento

1. Navigate to AllNow Assist Skill KitAgentic Evaluations.  
   You can also start from the testing page of the AI Agent Studio. Navigate to AllAI Agent StudioTesting. Select Start automated evaluation. You're redirected to the guided setup.
2. On the evaluations home page, select New evaluation run to begin the guided setup.
3. In the Add general info step, add a name and select the agentic workflow or AI agent that you want to evaluate.  


4. Select Continue to go to the next step.  
   Each time you navigate through a step, the evaluation run is saved automatically as a draft. At any point, you can select Save as draft.

   If you want to exit the guided setup, you can select Exit setup. You're redirected to the Agentic Evaluations page.
   * If you select Save and exit, the evaluation run appears on the Agentic Evaluations page with the status of Draft.
   * If you select Discard and exit, the evaluation run draft is deleted.
5. Select your evaluation metric.  
   Overall task completeness evaluation is selected by default. Running multiple evaluation metrics at a time can help provide a more comprehensive overview of the agentic workflow's or AI agent's performance.

   To see more information about each plan, you can expand the card for each evaluation plan by selecting the chevron icon .

   Any custom metrics that you have published appear as options as well. If you don't see your custom metric, make sure that it's published. See [Create a custom metric](https://servicenow-prod.fluidtopics.net/X~wQBquETgPw4HxvKp70SA "Create a custom metric for evaluating AI agents and agentic workflows to test the outputs against expected responses.") for more information.  
   Nota:  
   The tool calling correctness metric is not available for AI voice agents.


6. Choose your dataset.
   1. Select an existing dataset or create your own.
   2. If you're creating a dataset, give it a name and description.
   3. To create a dataset, choose whether you want to run the agentic workflow or AI agent to generate new execution logs or use existing ones.  
      Nota:  
      If you are evaluating AI voice agents, you must use existing execution logs.  
      {#execute-aia-eval__entry__2}

      | Field name | Description |
      |-|-|
      | Table | The source table for records that the agentic workflow or AI agent uses to perform tasks and create executions. |
      | No. of records | The maximum number of records within the dataset you want to run the evaluation on. If there are more records in the dataset than the maximum number of records, any records after the maximum number of records are ignored. |
      | Filters | Conditions for narrowing down the list of records for the agentic workflow or AI agent to use to generate execution log data. |
      | Starting phrase | Utterance given to the agentic workflow or AI agent to execute. Use the pill picker to select the specific dynamic inputs necessary to perform its task. For example, your starting instruction for evaluating an agentic workflow that generates resolution plans, you can set the starting instruction to <kbd class="ph userinput">Help me resolve {{incident.number}}</kbd>. Inputs from the record must be written between double curly braces. |
      | Additional business context | Information given to the Large Language Model (LLM) acting as the invoking user that supplements information found within the table records. For example, a tuition reimbursement agentic workflow must know the normal reimbursement allowance, so the additional business context could be a knowledge article with that information. |
      [Tabela 1. Configure dataset form for new execution logs]

      Nota:  
      If you are creating new execution logs, the user that submits the evaluation must successfully pass the ACLs of the AI agent or the agentic workflow and all of its AI agents. If the correct role requirements aren't met, the execution logs will only report that the user does not have access, and the evaluation will fail. See [Security for agentic AI](https://servicenow-prod.fluidtopics.net/22j7AP47_LhOTCRi_NkWGw "Implement security controls for AI agents and agentic workflows through access control lists (ACLs) and user identities to increase alignment with the access control-based security measures in the agentic system.") for more information.  
      {#execute-aia-eval__entry__14}

      | Field name | Description |
      |-|-|
      | No. of records | The maximum number of records within the dataset you want to run the evaluation on. If there are more records in the dataset than the maximum number of records, any records after the maximum number of records are ignored. |
      | Filters | Conditions for narrowing down the AI execution log records you want to include in the dataset. Nota: Filter conditions are not supported for creating datasets of AI voice agent execution logs. |
      [Tabela 2. Configure dataset form for existing execution logs]


   4. Select See preview to see a list of records based on the conditions you specified.  
      You can narrow down the records further by only selecting some of the records in the preview list. Unselected records won't be included in the dataset.
7. Review the agentic evaluation details in the last step of the guided setup.  
   If you want to make changes, you can select Back to go to a previous step, or you can select the step in the sidebar.


8. Select Start evaluation.

## Resultado

Your evaluation run executes. The time it takes for it to complete varies, but once completed you can select the evaluation from the Agentic Evaluations page to view the results.

For more information on the metrics on the results page, see [Agentic evaluation run results](https://servicenow-prod.fluidtopics.net/Wgy51VVans8hQ1O0oynXvg "Learn about agentic evaluation runs and the meaning behind different evaluation scores from the agentic evaluation results page.").

