---
sourceDocument: Yokohama Enable AI
sourceDocumentLink: https://servicenow-prod.fluidtopics.net/r/yokohama/intelligent-experiences

 Release :

    - yokohama

ft:locale :

    - en-US

ft:publication_title :

    - Yokohama Enable AI

ft:clusterId :

    - platai

bundleId :

    - platai

workflow :

    - Platform


---

# Evaluation tab

# Evaluation tab in AI Control Tower {#ariaid-title1}

Release version: Yokohama  
Updated July 8, 2025  
![](https://www.servicenow.com/docs/portal-asset/ico-clock) 4 minutes to read
Summarize  
![AI sparkle icon](https://servicenow.com/docs/portal-asset/ai-sparkle-icon) Summarized using AI  
This content was generated using new OpenAI-powered functionality. Results are provided on an as is basis and are not guaranteed to be accurate or complete.  

## Summary of Evaluation tab in AI Control Tower

The Evaluation tab in AI Control Tower provides a dedicated dashboard to measure, automate, and enhance the quality of Virtual Agent interactions.
It addresses key challenges in monitoring virtual agent conversations, aiming to improve the end-user experience and overall virtual agent performance.
This functionality is available for users with thesnaigovernance.aistewardrole and requires enabling evaluations.
Show full answer Show less  
Certain conversations are excluded from auto-evaluation, including HR-related conversations, those linked to inaccessible or empty Knowledge Base articles, immediate transfers to live agents without prior virtual agent interaction, short conversations below a configurable word count threshold, and any custom-defined exclusion triggers.

## Key Features

* **Automated Quality Measurement:** The dashboard automates the evaluation of conversation quality using large language models (LLMs) to sample and assess 10% of daily conversations, producing trustworthy metrics for performance tracking.
* **Comprehensive Metrics:** Conversations are evaluated across eight distinct metrics, each corresponding to a custom skill in the AI Skill Kit, such as coherence, intent accuracy, and smooth flowing conversation. These skills default to the Now LLM provider but can be switched to Azure OpenAI, Google Gemini, or AWS Claude for potential improvements.
* **Human Feedback Integration:** Users can provide manual feedback on conversations, which complements automated scores and helps benchmark and refine the evaluation process.
* **Insightful Dashboard Widgets:** The Overview tab displays average auto-evaluation and human feedback scores, evaluation trends, total evaluations, and detailed human feedback data to facilitate continuous monitoring and improvement.
* **Service Desk Manager Support:** Enables managers to track auto-evaluation scores over time and add manual feedback, providing actionable insights to improve conversation quality.
* **Sustainable and Scalable Process:** Combines automated evaluation with manual feedback to create a scalable system that evolves and sustains virtual agent performance improvements.

## Evaluation Process

Each day, 10% of conversations are sampled and assessed for eligibility based on exclusion criteria. Eligible conversation transcripts are sent to the configured LLM, which evaluates them against multiple metrics. Post-processing extracts scores and reasoning, storing results in the Evaluation tables. Note that evaluation timing affects score aggregation, as scores are assigned based on the evaluation date rather than the conversation creation date.

## Practical Considerations for ServiceNow Customers

* Ensure the **snaigovernance.aisteward** role is assigned to users managing evaluations.
* Configure exclusion criteria and thresholds to focus on meaningful virtual agent interactions.
* Leverage human feedback capabilities to supplement automated scores for richer analysis.
* Consider switching LLM providers within AI Skill Kit to optimize evaluation accuracy for your use case.
* Use dashboard insights to continuously monitor virtual agent performance and guide iterative improvements.
* Note that domain separation is not supported in the Evaluation dashboard.  
The Evaluation tab contains the Evaluation dashboard, which is designed to measure, automate, and improve the quality of interactions with Virtual Agent. This dashboard addresses several key challenges to enhance the end-user experience and overall virtual agent utility.

## Evaluation dashboard {#ai-evaluation__section_y5x_slr_hgc}

Prerequisites

Role required: sn_ai_governance.ai_steward

You must [Enabling evaluations](https://servicenow-prod.fluidtopics.net/3RN0IOJ_TNcy~D9XJgucHA "Evaluate random conversations by enabling continuous monitoring.").  
Conversations are excluded from auto-evaluation if any of the following conditions are met:

* HR conversations: Conversations related to Human Resources are filtered out, which means that they aren't evaluated.
* Inaccessible or empty Knowledge Base (KB) articles: Conversation involving a Genius Result that points to a KB article that is either not accessible via script or is empty. For example, certain restricted HR Knowledge articles.
* Immediate live agent transfer: A conversation that begins immediately with transfer to a live agent, with no prior interaction with the virtual agent.
* Short conversations: Conversations having fewer than 180 words before a live agent is invoked. The word count is configurable via the autoEvalConstants script Include. The assumption is that conversations below this threshold didn't contain a meaningful interaction with the Virtual Agent.
* Custom triggers: Any custom-defined exclusion triggers.
{#ai-evaluation__ul_q1v_kpr_hgc}

Evaluation dashboard overview  
The Evaluation dashboard helps in:

* Establishing a reliable measurement process by enabling the systematic tracking of the end-user experience with the Virtual Agent, providing deeper insights into interactions.
* Automation of conversation quality evaluation by automating the process of evaluating conversation quality across different user interactions. This automation helps lead to the creation of a trusted, scalable metric for performance tracking.
* Continuous improvement by supporting the iterative refinement of the virtual agent's performance, enhancing the overall user experience.
* Scalable monitoring by helping ensure that the process of evaluating and tracking virtual agent quality is both efficient and scalable, promoting quick identification of issues and improvements over time.
* User feedback integration through a set of optional questions enables you to provide direct feedback on their experience, which is used to improve the quality of future interactions.
* Service desk manager insights by enabling service desk managers to track and review auto-evaluation scores over time. Managers can also manually add feedback for benchmarking purposes, providing valuable insights into conversation quality and opportunities for improvement.
* Sustainable evaluation process by continuously improving virtual agent performance through a combined approach of automated evaluation and manual feedback enabling a scalable and sustainable system that evolves over time.
{#ai-evaluation__ul_ugx_wlr_hgc}  
Important:  
Evaluation dashboard doesn't support domain separation.

## Overview tab {#ai-evaluation__section_t1h_znn_xfc}

The Overview tab of the Evaluation dashboard provides a comprehensive view of all metrics and evaluation data.

The following widgets are available, showing various metrics:

* Average Auto Eval score for the selected metric: Shows the average auto-evaluation score for the metric selected and its trend over time.

  For more information about each metric, see [Evaluation metrics and calculations](https://servicenow-prod.fluidtopics.net/No0qLiWN~RBPvt1TbaAMyQ "Metrics against which conversations are evaluated and calculation of adjusted scores.").
* Average Human Feedback score for the selected metric: Shows the average human-labeled score for the selected metric.  
  Note:  
  The score is available only if there are sufficient chat records that are manually evaluated. For more information about manually evaluating conversations, see [Human feedback for evaluations](https://servicenow-prod.fluidtopics.net/Kbj1popW1Lz3MZqOdFQt8Q "Expand the Human feedback section to see details on evaluations and their satisfaction scores.").
* Evaluation score trend: Tracks the weekly score for the selected metric.

  If you turn on the View Deviation and Adjusted Scores toggle, it shows the comparison between the auto-evaluated and user-defined scores by overlaying the upper, lower deviations, and the final adjusted
  score on the trend chart.  
  Note:  
  The deviation and adjusted scores are calculated only if you have at least 50 human labels.

  For more information about how the calculations are made, see [Evaluation metrics and calculations](https://servicenow-prod.fluidtopics.net/No0qLiWN~RBPvt1TbaAMyQ "Metrics against which conversations are evaluated and calculation of adjusted scores.").
* Evaluations: Shows the total number of conversations that were evaluated each week.

* Human feedback section: Contains detailed information about each evaluation. From here, you can manually evaluate conversations. For more information, see [Human feedback for evaluations](https://servicenow-prod.fluidtopics.net/Kbj1popW1Lz3MZqOdFQt8Q "Expand the Human feedback section to see details on evaluations and their satisfaction scores.").
{#ai-evaluation__ul_ibl_1pl_rfc}

## Evaluations {#ai-evaluation__section_ndf_y2w_hgc}

Each conversation is evaluated on eight different metrics. For each of these metrics, there's a separate skill. You can view these skills in AI Skill Kit under Custom skills.

For more information about each metric, see [Evaluation metrics and calculations](https://servicenow-prod.fluidtopics.net/No0qLiWN~RBPvt1TbaAMyQ "Metrics against which conversations are evaluated and calculation of adjusted scores.").

Role required: sn_skill_builder.admin

The following Now Assist custom skills are used:

* Chat Topic Classifier
* Coherence Chat Evaluation
* Conciseness Chat Eval
* Context Retention
* Inadequate Slot Filling Chat Eval
* Intent Accuracy Chat Eval
* Smooth Flowing Conversation Chat Eval
* Truthfulness Hallucination Chat Eval
{#ai-evaluation__ul_qdm_jpg_ngc}

The default provider for these skills is Now LLM. You can change the provider to Azure OpenAI, Google Gemini or AWS Claude. Azure OpenAI has been observed to improve results in certain scenarios.

For more information about AI Skill Kit, see [AI Skill Kit](https://servicenow-prod.fluidtopics.net/dxWc7_6kndyvOpbI~VgDhA "Use ServiceNow AI Skill Kit to create and publish custom prompts and skills for Now Assist. Creating custom skills and prompts enables you to have greater flexibility with Now Assist's generative AI capabilities.").

Process of evaluation

Flow: Execute Evaluation.  
1. 10% of the daily conversations are sampled, checking if the conversation is good enough to be evaluated or not. The evaluation is done by building the transcripts for these conversations and then sending it to the set large language model (LLM).
2. For the conversations that are good enough to be evaluated, the transcripts along with the prompts for different scales are sent to the LLM and the LLM then evaluates the conversations.
3. After evaluation, the conversation goes through post processing, where the scores and the reason for scores that the LLM has provided are parsed and stored in the Evaluation and Evaluation Metrics tables.
{#ai-evaluation__ol_jqc_ndy_hgc}  
Note:  
Conversation evaluation estimates are considered as of evaluation date and not the conversation created date. For example, if a chat that happened at time t is evaluated at time t+10, the scores from evaluator is aggregated for the week of t+10 and not for the week of t.

For detailed information about the evaluation flow, see [Evaluation flow](https://servicenow-prod.fluidtopics.net/~GjYdu~lu74LijB2TsKsEQ "The workflow for evaluation execution, which performs evaluations when conversations are completed.").
**Related concepts**   

* [Value tab in the Evaluation dashboard](https://servicenow-prod.fluidtopics.net/RnEitTGAwt9b6zo1GvI8KQ "The Value tab in the Evaluation dashboard displays the value, efficiency, and time savings of the virtual agent. The information on this tab gives you a transparent and reliable estimation of the value delivered by the virtual agent, focusing on a quality-adjusted calculation of the time saved for users.")
* [Evaluation dashboard References](https://servicenow-prod.fluidtopics.net/TO7VoIWeycB2soa4XZ1Pxg "Reference topics for the Evaluation dashboard.")

