Evaluation Module – Comprehensive QA Testing Guide
1. Purpose & Overview
The Evaluation Module is a dedicated workspace for testing, measuring, and improving your bot’s performance.
It works hand-in-hand with datasets created in the Data Source Module (from CSV/Excel files).
By running evaluations, QA teams can identify weaknesses in accuracy, speed, and language understanding, then apply fixes and re-test.
2. Key Evaluation Metrics
Each metric targets a different aspect of bot quality:
| Metric | What It Measures | Why It Matters |
|---|---|---|
| Confidence | The certainty of the bot’s answers. | Reduces vague or low-confidence replies and flags when fallback handling is needed. |
| Relevance | How closely responses match the user’s intent. | Ensures conversations stay on topic. |
| Response Time | Time taken to generate a reply. | Improves real-time user experience by highlighting slow queries. |
| Classification | Correctness of categorical predictions (e.g., intent detection). | Directly impacts model accuracy and reliability. |
| Similarity | How close the bot’s answer is to the ideal reference answer. | Improves naturalness and semantic understanding. |
Similarity Methods
Select from multiple scoring techniques depending on your evaluation needs:
- Cosine Similarity
- Fuzzy Matching
- String Check Grader
- Euclidean Distance
3. Prerequisites
- Datasets: Must come from the Data Source module and be in CSV or Excel format.
- Permissions: Access to the bot space and Evaluation module.
- Clean Data: Ensure the dataset has clearly defined question–answer pairs for accurate scoring.
4. Navigation
- Login to the application.
- Click Bot Space in the left sidebar to enter the assessment area.
- Select Evaluation from the menu.
- The top section displays global stats such as Total Evaluations, Completed, Running, and Failed.
5. Creating a New Evaluation
Follow these steps to set up a new run:
- Start
- Click “+ Create New” in the top-right corner of the Evaluation dashboard.
- Name the Evaluation
- In the popup, provide a unique name (e.g., CustomerSupport_Eval_Q1).
- Select Dataset
- Choose a dataset previously prepared in the Data Source module (CSV/Excel).
- Choose Metrics
- Tick one or more metrics: Confidence, Relevance, Response Time, Classification, Similarity.
- If you include Similarity, select the scoring method (Cosine, Fuzzy, etc.).
- Run
- Click Run Evaluation.
- The new evaluation appears in the dashboard table with:
- Evaluation Name
- Dataset Name
- Criteria (metrics selected)
- Date & Time
- Status (Running, Completed, Failed)
- Actions (View, Edit, Download)
6. Updating an Existing Evaluation
- Locate the evaluation in the dashboard list.
- Click the Edit (pencil) icon in the Actions column.
- Modify:
- Name
- Dataset
- Selected metrics
- Click Update Evaluation to save changes and re-run if desired.
7. Viewing & Downloading Results
- Monitor status until it shows Completed.
- Click the Download icon to export results:
- Excel (XLSX) – full metric tables, ideal for deep analysis.
- CSV – lightweight format for further processing or integration.
- PDF – optional summary report.
- Review metrics:
- Overall Summary: Total/Completed/Running/Failed evaluations.
- Per-Metric Scores: Detailed breakdowns for each criterion.
Tip: Exporting in CSV or Excel makes it easy to feed evaluation results back into the Data Source module for retraining.
8. Best Practices for QA Testing
- Diverse Datasets: Use varied scenarios—different intents, edge cases, and user phrasings.
- Incremental Testing: Start with key metrics (e.g., Confidence, Relevance) before adding others.
- Continuous Improvement: Re-run evaluations after each model update to track performance trends.
- Failure Analysis: Investigate failed evaluations to uncover latency or classification issues.
- Version Control: Name evaluations with dates or version numbers (e.g., v1.2_March2025) for easy history tracking.
9. Example Workflow
- Prepare Data: Export approved Q&A pairs from Data Source as Excel.
- Create Evaluation: Select metrics Confidence, Relevance, and Similarity (Cosine).
- Run & Review: Download Excel results, filter low-confidence answers.
- Retrain Bot: Adjust bot logic or retrain on weak areas.
- Re-Evaluate: Re-run using the updated dataset to confirm improvements.
10. Summary
The Evaluation Module provides a structured, repeatable process to measure, analyze, and improve your bot’s real-world performance.
By leveraging CSV/Excel datasets, clear metrics, and downloadable reports, QA teams can maintain high-quality conversational AI with confidence.
Document Owner: Arjun Haldankar
Version: 1.0
Last Updated: 25-09-2025