Evaluation

Evaluation Module – Comprehensive QA Testing Guide

1. Purpose & Overview

The Evaluation Module is a dedicated workspace for testing, measuring, and improving your bot’s performance.
It works hand-in-hand with datasets created in the Data Source Module (from CSV/Excel files).
By running evaluations, QA teams can identify weaknesses in accuracy, speed, and language understanding, then apply fixes and re-test.


2. Key Evaluation Metrics

Each metric targets a different aspect of bot quality:

MetricWhat It MeasuresWhy It Matters
ConfidenceThe certainty of the bot’s answers.Reduces vague or low-confidence replies and flags when fallback handling is needed.
RelevanceHow closely responses match the user’s intent.Ensures conversations stay on topic.
Response TimeTime taken to generate a reply.Improves real-time user experience by highlighting slow queries.
ClassificationCorrectness of categorical predictions (e.g., intent detection).Directly impacts model accuracy and reliability.
SimilarityHow close the bot’s answer is to the ideal reference answer.Improves naturalness and semantic understanding.

Similarity Methods
Select from multiple scoring techniques depending on your evaluation needs:

  • Cosine Similarity
  • Fuzzy Matching
  • String Check Grader
  • Euclidean Distance

3. Prerequisites

  • Datasets: Must come from the Data Source module and be in CSV or Excel format.
  • Permissions: Access to the bot space and Evaluation module.
  • Clean Data: Ensure the dataset has clearly defined question–answer pairs for accurate scoring.

4. Navigation

  1. Login to the application.
  2. Click Bot Space in the left sidebar to enter the assessment area.
  3. Select Evaluation from the menu.
    • The top section displays global stats such as Total Evaluations, Completed, Running, and Failed.

5. Creating a New Evaluation

Follow these steps to set up a new run:

  1. Start
    • Click “+ Create New” in the top-right corner of the Evaluation dashboard.
  2. Name the Evaluation
    • In the popup, provide a unique name (e.g., CustomerSupport_Eval_Q1).
  3. Select Dataset
    • Choose a dataset previously prepared in the Data Source module (CSV/Excel).
  4. Choose Metrics
    • Tick one or more metrics: Confidence, Relevance, Response Time, Classification, Similarity.
    • If you include Similarity, select the scoring method (Cosine, Fuzzy, etc.).
  5. Run
    • Click Run Evaluation.
    • The new evaluation appears in the dashboard table with:
      • Evaluation Name
      • Dataset Name
      • Criteria (metrics selected)
      • Date & Time
      • Status (Running, Completed, Failed)
      • Actions (View, Edit, Download)

6. Updating an Existing Evaluation

  1. Locate the evaluation in the dashboard list.
  2. Click the Edit (pencil) icon in the Actions column.
  3. Modify:
    • Name
    • Dataset
    • Selected metrics
  4. Click Update Evaluation to save changes and re-run if desired.

7. Viewing & Downloading Results

  1. Monitor status until it shows Completed.
  2. Click the Download icon to export results:
    • Excel (XLSX) – full metric tables, ideal for deep analysis.
    • CSV – lightweight format for further processing or integration.
    • PDF – optional summary report.
  3. Review metrics:
    • Overall Summary: Total/Completed/Running/Failed evaluations.
    • Per-Metric Scores: Detailed breakdowns for each criterion.

Tip: Exporting in CSV or Excel makes it easy to feed evaluation results back into the Data Source module for retraining.


8. Best Practices for QA Testing

  • Diverse Datasets: Use varied scenarios—different intents, edge cases, and user phrasings.
  • Incremental Testing: Start with key metrics (e.g., Confidence, Relevance) before adding others.
  • Continuous Improvement: Re-run evaluations after each model update to track performance trends.
  • Failure Analysis: Investigate failed evaluations to uncover latency or classification issues.
  • Version Control: Name evaluations with dates or version numbers (e.g., v1.2_March2025) for easy history tracking.

9. Example Workflow

  1. Prepare Data: Export approved Q&A pairs from Data Source as Excel.
  2. Create Evaluation: Select metrics Confidence, Relevance, and Similarity (Cosine).
  3. Run & Review: Download Excel results, filter low-confidence answers.
  4. Retrain Bot: Adjust bot logic or retrain on weak areas.
  5. Re-Evaluate: Re-run using the updated dataset to confirm improvements.

10. Summary

The Evaluation Module provides a structured, repeatable process to measure, analyze, and improve your bot’s real-world performance.
By leveraging CSV/Excel datasets, clear metrics, and downloadable reports, QA teams can maintain high-quality conversational AI with confidence.


Document Owner: Arjun Haldankar
Version: 1.0
Last Updated: 25-09-2025

Next

Leave a Reply

Your email address will not be published. Required fields are marked *