Overview
When a scorecard is marked Automated, AutoQA answers its points for you. Scorecard accuracy improvement is the workflow that tells you how often AutoQA agrees with your human reviewers, and gives you a structured way to make it agree more often.
The workflow is built around two ideas:
A point carries your business criterion — the name and description you wrote.
A definition version is one particular AI instruction for applying that criterion. A point can accumulate many versions over time, but only one is Active — the one AutoQA is executing right now.
Improving accuracy means finding a better version and activating it, without ever changing the business wording your reviewers read.
Everything below happens on an automated scorecard. For building and managing the scorecards themselves — categories, points, partial scoring, archiving, import and export — see Scorecards.
The improvement workflow is labelled Beta.
How It Works
Where to find it
Open Settings from the main navigation, click Scorecards, and open an automated scorecard.
The settings sidebar shows Scorecard Accuracy — the weighted accuracy across the whole scorecard, based on human-validated reviews.
Each point row shows its own accuracy value. Click that value to open the Accuracy details dialog, where the whole improvement workflow lives.
If a point has neither observations nor an automated scorecard behind it, no accuracy element is shown at all.
Who can use it
Reading accuracy is available to anyone who can open the scorecard. Every action that changes what AutoQA executes — Improve, Activate, Delete, Edit definition, and the Evaluate actions — requires the team-management permission, the same permission needed to save a scorecard. The full access table is in Scorecards.
Reading the accuracy value
Point accuracy is measured per definition version, and counts only the human-validated AutoQA scores produced by the version currently in use. Until that version has five of them, the row shows Not enough data instead of a percentage, with a tooltip explaining that at least 5 validations are needed. Hovering the value gives guidance text, the same way the scorecard-level indicator does.
Because the count follows the active version, activating an improved candidate — or editing the point's description, which also creates a version — sends the row back to Not enough data until five conversations have been scored by the new version and reviewed. Nothing is lost: the earlier measurements stay on their own version in the accuracy chart.
The Accuracy details dialog contains:
a bar chart of accuracy per version, with confidence bars and the version currently in use marked (Active). Only versions that already have reviewed scores get a bar, so a version you have just activated is absent from the chart — including the (Active) marker — until its first review arrives
the point's Versions list
the improvement steps and the Advanced panel
If the scorecard was switched back to manual but still has past AutoQA observations, the value stays clickable, but the dialog then shows only the version history and definition details — the accuracy chart and the improvement workflow are hidden.
Step 1: build the evidence by reviewing conversations
Improve accuracy (by following these steps) walks you through three steps:
Create score point — done as soon as the point exists.
Review conversations — while the point is not ready yet, click this step (or the Review conversations button under the steps) to jump straight to the Conversations screen in review mode, with this scorecard and the AI's own scores already selected in the review panel, so you can confirm or correct them. The conversation list itself keeps whatever filters you normally use — the shortcut does not narrow it down to conversations this scorecard scored, so set that filter yourself if you want to work through them.
Get improvements — stays disabled until enough reviews exist.
Two chips track your progress towards step 3: Yes n/5 and No n/5, with the hint "Review at least 5 Yes and 5 No conversations to get improvements." Only clean Yes and No answers count: a review scored 0 counts as No, a review scored at the point's maximum counts as Yes, and a review marked N/A is collected as an N/A example. Partial scores still count towards production accuracy but are not used as improvement examples. N/A examples are used when available but never block readiness.
Once both counts are reached, the hint, the chips and the two review shortcuts are replaced by the Improve button. You can keep reviewing conversations from the Conversations screen at any time; the shortcuts simply stop being offered here.
Eligible Yes, No, and N/A reviews form one reusable set for the scorecard point, following the same principle as Topics: a review describes the business criterion, while definition versions are different AI instructions for applying that criterion. The set is reused whenever you improve, edit, or reactivate a version of that point.
Step 2: run an improvement
Reaching the requirement enables Improve; nothing starts automatically. Each cycle is triggered manually, and further reviews keep enriching the same reusable set. While the run is in progress the dialog shows Improvement queued and then Improvement running with a progress bar.
There is no success message. When the run finishes, the Get improvements step is simply marked as done and the candidate appears in the Versions list, from where you activate it.
A run that found nothing worth changing leaves No validation mistakes found on the version instead of creating a candidate.
A failure is shown as Improvement failed with the error text and a Retry button.
While an improvement or an evaluation is running, the point's description field in the scorecard form is locked, with the hint "Available once the current run finishes". The same applies to Edit definition, Activate and Delete.
Step 3: decide what to activate
A candidate is never activated automatically. Review its instruction and its evaluation result, then click Activate if you want AutoQA to start using it.
The compact Improvement ready cue appears when at least one retained candidate has an evaluated accuracy at least five percentage points above the active version's latest evaluated accuracy. It is a point-level cue measured against the best of the retained candidates, not against the newest one, so on a point that has collected several candidates it does not tell you which one qualifies — an older candidate can keep the cue visible while the latest one does not clear the threshold. Open Versions and compare the evaluation results there before deciding what to activate. When either evaluation is missing, the candidate still appears in the version history, just without that recommendation. These rolling evaluations can contain different reviewed conversations, so the cue is a selection aid rather than a paired experiment.
If you activate a different version while a candidate is still waiting for your decision, the dialog warns that the candidate was produced from a definition that is no longer active and asks you to review it carefully before activating it. Reviewing more conversations does not raise this warning — it simply enriches the set for the next improvement run.
Versions
The Versions card lists every definition version of the point, newest first. Each row can show:
its creation date followed by where it came from — Original description, Manual, or Improvement
its accuracy, or Not enough data — only while the scorecard is automated; on a scorecard switched back to manual the rows carry no accuracy at all
Active for the version AutoQA is currently using
Instruction pending or Evaluating… while work is in progress
Improvement queued, Improvement running, No validation mistakes found or Improvement failed — the status of the latest improvement run attached to that version. On an original or manual version this is the run you started from it, not a problem with the version itself. Successful runs are not shown; they surface as a new candidate row instead.
Hovering a row reveals:
Activate — switch AutoQA to that version. An improvement-sourced version becomes activatable only after its improvement run and its candidate evaluation have both finished successfully.
Delete — available for automated points on inactive versions only. You are asked to confirm ("Delete definition version #n? This cannot be undone."). Deleting removes the version from the list without erasing which version scored past conversations.
Because AutoQA records the exact version it used, later description edits or reactivating an older version never move historical results between versions. A retained version can be activated again later provided it has an AI instruction and, if it came from an improvement, that improvement and its candidate evaluation both finished successfully. A candidate left behind by a failed run stays visible in the history for reference but cannot be activated.
The instruction the AI uses
Each version stores an Instruction the AI uses — the compact text AutoQA actually executes. The initial point description and each later description change create an instruction version built from that description exactly; an empty description stays an empty AI instruction. Renaming a point updates its label without creating a new version.
A version created this way is listed as Manual and starts being used by AutoQA as soon as the scorecard is saved — there is no separate activation step, and the point's accuracy restarts from Not enough data. This is the opposite of Edit definition in Advanced, which stores an inactive version that you activate yourself. If you want to try new wording without changing what AutoQA does today, use Edit definition rather than the description field.
The name and description in the scorecard form always remain your text. Activating an AI-improved version changes only the instruction AutoQA executes, never what you typed.
When an instruction is still being generated
When a version's instruction is still being generated, the dialog shows an Instruction pending notice whose only action is Generate AI instruction. The version is not frozen otherwise: Edit definition in Advanced and Delete in Versions stay available, so you can revise or discard it while the instruction is still missing. What the missing instruction blocks is activating, evaluating and improving that version.
The dialog checks for the result every 5 seconds, but only from the moment you saved the version or pressed Generate AI instruction. If you close the scorecard or reload the page and come back later, the notice still shows the pending state while nothing is being checked — press Generate AI instruction to start checking again.
If generation makes no progress for about ten minutes of checking, the dialog stops and the notice becomes a warning that the instruction did not arrive. Two recovery actions appear at that point, and Generate AI instruction becomes Retry generation:
Check status — ask again whether the instruction is ready, in case it arrived after the dialog gave up.
Enter manually — write the instruction yourself instead of waiting for another attempt. This opens the full definition editor with Edit AI instruction already ticked, and saving creates a new inactive version carrying your instruction. The timed-out version stays in the history as Instruction pending, so activate the new version once you are happy with it.
A version with no instruction yet cannot be activated and cannot be used as the starting point for an improvement or an evaluation.
Advanced: editing and testing a definition yourself
The Advanced panel inside Accuracy details exposes the full definition behind the selected version:
Criterion — the structured rule behind the version. AutoQA does not send this text to the AI when it scores conversations; it executes the Instruction the AI uses that was generated from the criterion.
Score Yes when / Score No when / Score N/A when — the explicit conditions for each outcome.
Boundary cases — tricky examples tagged Yes, No, or N/A with the reason.
Change summary — what changed compared with the previous version.
Description this version was built from — a historical reference, shown only when the description that produced the version differs from the point's current description.
Evaluate — test the production AI instruction against reviewed conversations. Open its dropdown and choose Evaluate structured definition to test the full definition instead. The last successful result is shown as accuracy, sample count, and finish time, with the two kinds of result reported separately. A run that failed shows an Evaluation failed message with a Retry button instead of those numbers.
The button appears once the point has at least one reviewed example on an automated scorecard you can edit. The main Evaluate action needs the same 5 Yes and 5 No reviews that unlock Improve, so before that it stays greyed out while Evaluate structured definition in the dropdown is already usable — provided the selected version has a structured definition behind it. Both are unavailable while an evaluation is already running, while the version's AI instruction is still pending, and while you have the Full definition dialog open in edit mode.
Edit definition — open the Full definition dialog to change the criterion, the Yes/No/N/A conditions, the boundary cases, and (with the Edit AI instruction checkbox) the instruction itself. Save as new version stores it as a new inactive version, so nothing changes for AutoQA until you activate it. If you leave Edit AI instruction unchecked, the new version starts with no instruction and sits in Instruction pending while one is generated from your edited criterion.
Evaluation and improvement use frozen inputs and can run at the same time, but a second run of the same kind is blocked until the first one finishes.
How accuracy is measured
The version history distinguishes two measurements, and they answer different questions:
Production accuracy compares each AutoQA answer with the human answer that replaced it. It counts exact matches for all score values, including partial scores, and becomes reportable after five reviewed results. Use it to understand live behaviour.
Candidate evaluation tests an AI-improved candidate on a separate subset of reviewed conversations, kept apart from the ones the candidate was written from. When that subset is too small, or has no Yes and No examples to test against, it is topped up with conversations the candidate was written from — so on a point with few reviews the candidate's score can look better than it will be in production. The more conversations you review, the less this happens. The active definition is not re-evaluated on the same subset either, so this result is not a direct base-to-candidate comparison. Use it as supporting evidence when deciding whether to activate a candidate.
Limits
Archived and template-managed scorecards cannot be improved, and the improvement and evaluation controls are unavailable on them. You can still open Accuracy details and read the chart, the versions, and the full definitions.
Tuning evidence is specific to the workspace conversations it was collected on, so it does not travel with a scorecard. A copied or spreadsheet-imported scorecard starts its accuracy history from scratch — see Exporting a scorecard in the Scorecards article for exactly what each route carries.
Getting better results
Most low automated accuracy comes from the scorecard, not from the AI:
Write point names and descriptions that produce a clear Yes/No answer, where Yes is the positive assessment and No is the negative one. Reword a question if it cannot be answered Yes/No — for example, change "Conversation sentiment" to "The sentiment of the conversation was positive".
Avoid "if applicable" wording. Instead, route only the conversations that are actually in scope to the scorecard, using filters on topic, direction, team, or queue.
Avoid subjective or overlapping questions, and questions about data held in third-party systems.
Review conversations steadily rather than in one burst. The reviewed set is what both the improvement runs and the evaluations draw on, so more reviews make every later decision more reliable.
