Overview
A scorecard is the structured rubric Ender Turing uses to evaluate a conversation. It groups your quality criteria into categories, and each category contains one or more points (also called scoring parameters) that an evaluator — a human reviewer, or the AutoQA engine — answers when scoring a conversation.
The Scorecards screen is where you build and maintain those rubrics for your workspace:
create a new scorecard from scratch, copy an existing one, or import one from an Excel file
edit categories, points, scores, and behaviour settings
mark a scorecard as the workspace default
mark a scorecard as automated so it can be applied by AutoQA
export a scorecard to Excel
archive scorecards you no longer want active, and restore them later
Scorecards live separately from teams. Linking a scorecard to a team and choosing each team's default scorecard happens from the Users & Agents area, not from the Scorecards screen itself.
How It Works
Where to find Scorecards
Open Settings from the main navigation, then click Scorecards.
The page opens on the list of scorecards in your workspace. A top-right toggle switches between Active and Archived views.
Who can use Scorecards
Only users whose role includes the team-management permission see the Scorecards entry in the Settings menu and can open the list. Users without that permission do not see the menu item and cannot reach the page.
Within the page, access splits into two levels:
Action | Required access |
Open the Scorecards list, see scorecard names, teams, accuracy, labels | Team-management permission |
Create, import, archive, unarchive, or delete a scorecard | Team-management permission |
Save changes to a scorecard's categories, points, or settings | Team-management permission |
Open a scorecard's edit page from a direct link | Scorecard-management permission |
Export a scorecard to Excel | Scorecard-management or team-management permission |
A user with only the scorecard-management permission can reach a scorecard's edit page from a direct link and inspect its structure, but saving any change still requires the team-management permission — the Save call will be rejected otherwise. The standard Settings → Scorecards navigation entry only appears for users who have the team-management permission.
A scorecard that you cannot edit because it is managed by an external Template appears with a read-only indicator and shows a "view" icon instead of an edit pencil. You can still open it to inspect its structure.
The Scorecards list
Each row shows:
Name — click it to open the scorecard. The row may also display a Default Scorecard chip if the scorecard is the workspace default, and an Automated chip if AutoQA can use it.
Conversations — how many conversations have already been scored against this scorecard.
Accuracy — for automated scorecards only, the weighted accuracy based on human-validated reviews. While the value is being computed you see a spinner; once available, the icon and colour signal whether AutoQA accuracy on this scorecard is in a healthy range or needs tuning. Hover the value for guidance text.
Teams — the teams this scorecard is linked to. If the scorecard is linked to more than one team, a +N chip lists the additional teams in a tooltip.
Labels — coloured labels you can attach for quick filtering.
A quick filter at the top of the page lets you narrow the list by:
a text search on the scorecard name
one or more teams (including a "no team" group for scorecards that are not linked to any team)
one or more labels
Each row's action area offers:
a Pencil (edit) or View icon to open the scorecard
an Archive / Unarchive icon
a Delete icon — shown only for archived, non-read-only scorecards
a Template info icon — shown only when a templates-library extension is connected to your workspace
Creating a scorecard
From the list, click + Create card in the top bar. A dialog appears with two tabs:
Create from blank — start a new scorecard with one empty Main Category and no points. Pick the Scorecard option in this tab to begin.
Create from existing — pick an existing scorecard to copy. The new scorecard is named Copy of \<original name\> (with a numeric suffix if a copy already exists) and Ender Turing immediately opens it for editing.
After creating, you land on the scorecard's edit page.
Importing a scorecard from Excel
Click + Import card in the top bar. Choose an .xlsx file and click Upload. The platform parses the file and creates a scorecard with the categories and points described in the spreadsheet. If the file is in the wrong format, or a scorecard with the same name already exists, an inline error explains the problem and no scorecard is created.
After a successful import, Ender Turing opens the new scorecard so you can fine-tune it.
Editing a scorecard
The edit page has two columns.
The left column contains the scorecard's categories, each shown as an expandable panel:
Click a category title to expand or collapse it.
Use the drag handle on the left of a category to reorder it within the scorecard.
Inside the category, each row is a point (scoring parameter) with these editable fields:
Point name and an optional description that explains what the reviewer (or AutoQA) is looking for.
Allow partial scoring — a checkbox. When off, the point is a Yes/No question and the score is its Point score (max) value. When on, you can pick one or more partial values that the reviewer can choose between (for example 0, 3, 5).
Point score — the maximum score for the point (used to populate the Yes/No score, or as the upper bound of partial values).
Critical — a switch (available for standard score-type scorecards) that flags the point as critical. If the critical point is failed, the whole category is considered failed.
A category drag-handle, an Edit name pencil, and a Delete icon are available on hover.
Use + Add parameter at the bottom of a category to add another point.
Use + Add category under the list to create a new category.
The right column is the scorecard's settings sidebar:
Scorecard Accuracy (shown only when the scorecard is marked Automated). The weighted accuracy from human-validated reviews and the explanatory tooltip help you decide whether the AI-driven scoring needs more validation or more descriptive points.
N/A buttons behaviour — what happens when a reviewer marks a point as N/A:
Exclude "N/A" from calculations
Count "N/A" as Max score
Count "N/A" as 0
Critical points behaviour — what a failed critical point does to the total score:
Include in total score calculations (Default)
Exclude critical points from calculations (only resets to zero)
Teams assigned — a read-only list of the teams this scorecard is linked to. Linking is managed from the Users & Agents area; the scorecard editor only displays the current assignment.
The top bar of the edit page contains:
a back arrow to return to the list
the scorecard name (editable, required)
Set as Default — toggle this on to mark the scorecard as the workspace default. Once it is the default, the toggle is replaced by a Default Scorecard chip.
Use as Automated — toggle this on to mark the scorecard as automated so AutoQA can apply it. Off by default for new scorecards.
an Excel export icon that downloads the scorecard structure as an
.xlsxfilea Cancel button (asks for confirmation, then reverts pending changes)
a Save button (disabled while validation is failing or in progress)
When you save changes that would alter historical scoring results — adding a new point, removing a point, changing a point's max score, changing its partial-score values, flipping its critical flag, or changing the N/A buttons behaviour — Ender Turing asks you to confirm before proceeding. The confirmation explains that historical scoring results may change significantly. Other edits (renaming a point, updating descriptions, reordering, changing partial-scoring eligibility without changing values) save without that warning.
If you try to leave the page with unsaved changes, you are asked to confirm that it is safe to discard them.
Removing points safely
When you save a scorecard, Ender Turing checks the points you are removing against any Enders (automations) that reference them. If at least one Ender — active or archived — uses one of the points you are removing, the save is blocked and the platform lists which points must stay. This is stricter than archiving the whole scorecard, because an archived Ender can be re-enabled later and would still expect those point IDs to exist.
While editing, attempts to delete points whose IDs are referenced elsewhere open a Unable to remove current point(s) dialog that lists the dependent entities (such as Enders) point by point, so you can remove or unwire those references first.
Marking a scorecard as Default
Toggling Set as Default on a scorecard makes it the workspace-wide default scorecard that reviewers see first when they open the scoring panel. Only one scorecard can be the default at a time; turning a different scorecard into the default replaces the previous one.
Note that a workspace default is separate from the per-team default scorecard that teams can have configured in the Users & Agents area. The per-team setting takes priority when a team is set, while the workspace default acts as a fallback.
Marking a scorecard as Automated
Toggling Use as Automated lets AutoQA score conversations against this scorecard. After at least five AI/LLM auto-scored conversations have been human-validated, the Scorecard Accuracy indicator shows the weighted accuracy and gives actionable guidance text (for example, "The AI/LLM instructions need improvement" or "Validated accuracy is reasonably high").
Accuracy per point
On an automated scorecard, each point row also shows its own accuracy value. It is measured per instruction version and counts only the human-validated AutoQA scores produced by the version currently in use. Until that version has five of them, the row shows Not enough data instead of a percentage, with a tooltip explaining that at least 5 validations are needed. Hovering the value gives the same kind of guidance text as the scorecard-level indicator.
Click the value to open the Accuracy details dialog, which holds the accuracy chart, the point's definition versions, and the tools for improving what AutoQA executes. That workflow is covered in its own article: Scorecard Accuracy Improvement.
If the point has neither observations nor an automated scorecard behind it, no accuracy element is shown at all.
Editing a point on an automated scorecard
Each definition version stores an Instruction the AI uses — the compact text AutoQA actually executes. Editing a point's description and saving the scorecard creates a new version built from that text and puts it straight into use, which also resets the point's accuracy to Not enough data. Renaming a point updates its label without creating a new version.
The name and description in the scorecard form always remain your text — activating an AI-improved version changes only the instruction AutoQA executes, never what you typed. To try new wording without changing what AutoQA does today, use Edit definition inside Accuracy details instead, which saves an inactive version. See Scorecard Accuracy Improvement.
Archiving and unarchiving
The list's Archive action on each row moves a scorecard out of the active list. Archived scorecards still exist (their historical scores are preserved) but they are hidden from scoring pickers and from the active scorecard list.
Archiving a scorecard that is used in active Enders is blocked. The archive icon is disabled and a tooltip explains that you must remove the Ender's usage of the scorecard first.
Archiving a scorecard that is only referenced by archived Enders is allowed.
The first time you archive a scorecard, Ender Turing asks for confirmation before proceeding. Unarchiving from the Archived view does not need confirmation.
Switch between Active and Archived in the top-right toggle to view either set.
Deleting
Deleting a scorecard is permanent and is intentionally restricted:
The scorecard must already be archived (the delete icon does not appear in the Active view).
The scorecard must not be linked to any team.
No Enders or scored conversations may depend on its points.
If teams or dependent entities still reference the scorecard, the delete icon is disabled with a tooltip listing what needs to be removed first (for example, "team", "conversations", "Enders"). When all dependencies are clear, click the Delete icon and confirm in the dialog to remove the scorecard.
Exporting a scorecard
On the edit page's top bar, the Excel-export icon downloads the scorecard structure — categories, points, score values, critical flags, and descriptions — as an .xlsx file. You can edit that file and re-import it as a new scorecard if you want to clone a scorecard between workspaces or share the rubric for offline review.
Copying an automated scorecard inside the workspace carries its current point text and active runtime instructions, but not its reviewed examples, accuracy observations, candidate evaluations, or version history. Spreadsheet export/import carries the point text and creates a fresh simple definition whose rule and runtime instruction use that text. Tuning evidence is specific to the workspace conversations on which it was collected.
Labels and organization
Each scorecard in the list has a Labels cell where you can attach coloured labels. Click the + to add an existing label or create a new one. Once attached, you can filter the list by one or more labels using the quick filter at the top of the page.
Tips for AutoQA scorecards
Automated scoring depends heavily on how unambiguous the scorecard is:
Each point should be answerable with a clear Yes = positive assessment, No = negative assessment. Reword a question if it can't be answered Yes/No (for example, change "Conversation sentiment" to "The sentiment of the conversation was positive").
Avoid "if applicable" wording. Instead, route only the conversations that are actually in scope to the scorecard, using filters on topic, direction, team, or queue.
Use the Critical flag for behaviours that should immediately fail the category, not for ordinary questions.
Validate at least five AutoQA-scored conversations per point so the platform can compute meaningful accuracy and tell you whether the point needs better descriptions or more validations.
