Skip to main content

Topics

Group conversations into reusable themes with generative AI, filter rules, or human labeling; tune Gen AI topics by reviewing samples; use topics in filters, dashboards, charts, and reports. This article is automatically maintained (topics).

Overview

Topics group conversations into reusable themes so that you can analyze and filter by what was actually discussed — for example, Refund request, ATM issue, or Verification complete. Each topic carries a name, an optional description, a categorization type (Use Gen AI, Use Rules, or Human Only), labels for organization, and a priority position used when a conversation matches more than one topic.

Topics show up everywhere conversations are explored: as filter chips in Conversations, breakdown buckets on the Dashboard, comparison cohorts in Charts, steps in Funnels, and pickers for human reviewers on a conversation. This article covers the Topics configuration page, where the topic catalog is created and maintained, and the single-page workbench for tuning Gen AI topics with reviewed historical samples.

How It Works

Accessing Topics

Open Analytics configuration from the main navigation, then select Topics. The page lists every topic in the workspace and opens with the Active view selected by default.

Who can use Topics

The Topics menu entry is only visible to users with the analytics-management permission. These users can create, edit, archive, restore, delete, import, and export topics, reorder topics or change a topic's position, and use the workbench on Gen AI topics — including the launched review mode and the version history.

Users without that permission do not see the Topics entry in the navigation and cannot reach this configuration page. Existing active topics still apply to their conversations and remain available wherever topic-based filters and analytics are shown — those users simply cannot change the topic catalog or tune its Gen AI definitions.

Choosing a categorization type

Each topic is one of three types. The type is picked when the topic is created and can be changed later by editing the topic.

  • Use Gen AI — a generative-AI model decides whether a topic applies based on a description you write. Best for new or evolving themes where you cannot pin down exact phrases. Requires a generative-AI provider to be configured for the workspace (see System). When a provider is configured, Use Gen AI is the default type for new topics. When no provider is configured, the Use Gen AI option is hidden and the Human Only type is used as the default instead. Gen AI topics unlock the workbench — the Versions rail, the Use this version promotion offer, the Review queue, and the Finish review action — on the topic page.

  • Use Rules — the topic applies when a filter expression matches the conversation. Best when the topic has clear, repeatable signals such as specific tags, words, or metadata. Use Rules is a legacy type: it is hidden from the type picker when you create a new topic and only appears as an option when you edit a topic that was already saved as rule-based. When a rule-based topic is created or its rules are edited, existing auto-assignments for that topic are cleared and conversations from the last 30 days are rematched against the new rules. Manually verified assignments are preserved; auto-assignments on conversations older than 30 days are not restored unless rematched manually.

  • Human Only — the topic has no matching logic of its own, so the platform never matches it automatically the way Gen AI and rule-based topics do. It shows up as an option for reviewers who label a conversation manually, and an Ender can still apply it by selecting the topic directly (see Which Enders use a topic). Best for internal labels, training cohorts, or anything that needs a human decision.

How many topics a conversation gets

Use Gen AI topics are applied by an Ender (automation), and the Ender decides whether a conversation can carry several Gen AI topics at once or just one. Rule-based and Human Only topics are not affected by this choice.

Open Enders from the main navigation, edit the Ender that sets topics, and look at its Set topics action. As soon as the action's topic selection includes any Gen AI topic — directly or through a bulk scope such as All Gen AI topics — the note "Some of selected topics are Gen AI, select a prompt" appears together with a Prompt to use dropdown. Two options are offered:

Prompt to use

What the AI assigns

Default Categorisation (multi-topic)

Every topic that is clearly justified by the transcript. A conversation can end up with several topics.

Single-topic Categorisation

At most one topic — the primary and dominant purpose of the conversation. Topics that are only secondary, brief, or incidental are not assigned.

Nothing is preselected, and the Ender cannot be saved with a Gen AI topic selected until you pick a prompt. Each entry in the dropdown shows its version and creation date. Archived prompts are still listed for reference but are marked with an Archived chip and cannot be selected.

Pick Single-topic Categorisation when a conversation should never carry more than one topic — for example a handful of deliberately close topics where one of them should always win. The cap works per action: it assigns at most one topic out of the topics that Set topics action selected, and it does not guarantee a topic is assigned at all. A second Ender whose action selects different topics can still add one of its own, and rule-based and Human Only assignments are independent of it. If you need one topic per conversation, keep the whole set in a single action. Pick Default Categorisation (multi-topic) when a conversation can legitimately be about more than one thing.

When no topic is assigned

With either prompt, a conversation that the AI treats as a misdial or an accidental contact with no valid request is left without a topic. With Single-topic Categorisation the same happens whenever none of the configured topics matches.

This is a normal outcome, not an error — the conversation finishes processing as usual and simply carries no topic. In topic breakdowns and filters those conversations are grouped under No topic. If you see more of them than you expect, the usual causes are topic descriptions that are too narrow or a topic catalog that does not cover the traffic.

Designing your topic catalog

If you are moving to Ender Turing from a manual topic catalog — the kind agents click through after a call, usually a tree three or four levels deep — the catalog needs rework before it will classify well here. A tree exists to keep a human from choosing between hundreds of options at once. That constraint does not apply to automatic classification, and carrying the tree over unchanged is the most common reason topic accuracy disappoints.

The catalog is flat. Topics have no parent/child relationship and cannot be nested. Each topic is one entry in a single list with a Position and optional Labels you can use to group topics for your own navigation. Use Labels for the grouping a tree used to provide; do not try to encode a hierarchy in topic names.

Position is a reporting rank, not a matching rule. Position does not influence which topics the AI assigns — it is never shown to the model. It is applied afterwards, to pick which of the topics already assigned to a conversation counts as its main one in analytics views such as Only main topic. Reordering two overlapping topics therefore changes which one is reported as primary, not whether either is matched. To change what actually matches, edit the topic descriptions (see Definition details and the slide-over drawer).

You do not have to force one topic per conversation. A conversation that covers a failed payment, a return request, and a courier complaint is genuinely about three things. Under Default Categorisation (multi-topic) it can carry all three, so each one reaches your statistics instead of only the branch an agent happened to pick. Single-topic Categorisation assigns at most one topic out of the topics its own Set topics action selected — it is not a guarantee that every conversation ends up with exactly one. A conversation that matches nothing in the catalog still lands under No topic (see When no topic is assigned), and a second Ender selecting a different set of topics can add one on top. If your reporting needs every conversation in exactly one bucket, the catalog has to cover your traffic and all the competing topics have to sit in one action; the prompt choice alone will not do it.

Put the meaning in the description, not just the name. Both the topic name and its description are sent to the AI. A name on its own — Error, Failure, Other issue — carries very little signal about which real conversations belong to it. Write a description that explains what the topic covers, the situations that count, and the wording customers actually use for it. For Gen AI topics you can go further and pin down the edges in the structured definition, using Include when, Exclude when, and Boundary cases (see Definition details and the slide-over drawer).

Separate topics that overlap in meaning. Nothing stops you from creating duplicates. Creating topics one at a time — through + Create or the API — applies no name check at all, so two topics both named Refund can sit side by side in the catalog. Only + Import rejects conflicts, and only against the catalog you already have: it compares each name in the file case-insensitively against every existing topic, archived ones included, but it does not compare the file's own rows against each other — a spreadsheet listing both Refund and refund creates both (see Importing and exporting). And no path detects two topics that mean the same thing under different names — a Refund topic and a Money back topic will both be offered to the AI with no warning. Legacy trees accumulate these because the same meaning was repeated under different branches, so review the catalog yourself: merge them, or if both must exist, use Exclude when and Boundary cases on each to state which one wins.

Prune before you scale up. There is no cap on how many topics a catalog can hold. Whether an unused topic costs you anything depends on how your Enders select topics: an Ender configured with a bulk scope such as All Gen AI topics evaluates every active topic of that type against every conversation it runs on, so narrow topics that almost never occur add noise and near-duplicate pressure. An Ender that lists topics individually only evaluates the ones you picked. Either way, archive topics tied to finished campaigns, retired products, or one-off issues (see Archiving and restoring).

Archiving is not free of reporting consequences, though. The assignments themselves are kept and come back intact when you restore the topic, but while it is archived it drops out of topic analytics — dashboards, charts, and reports skip archived topics, so conversations that carried only that topic are counted under No topic for the period. Archive when you no longer want the topic in your reporting; keep it active if you still report on the campaign it covers.

A new topic does nothing until an Ender uses it. Creating a topic of any type adds it to the catalog; it is only classified against conversations once it is included in an Ender's Set topics action. Gen AI and rule-based topics can be included either directly or through a bulk scope such as All Gen AI topics; a Human Only topic has to be picked individually, because the bulk scopes cover only the Gen AI and rule-based types. Saving rules on a rule-based topic is not a substitute for an Ender: it rematches conversations from the last 30 days once (see Choosing a categorization type), but new conversations are still only tagged by an Ender that selects the topic.

After migrating a catalog, use the Used by Enders column to find the gaps, but read it carefully — it is reliable in one direction only. Not used means nothing references the topic at all, which is the signal you want. A non-zero count does not prove the topic is being classified: the count also includes Enders that merely reference the topic in a trigger filter, and Enders that are inactive or archived. To confirm a topic is really live, open it and check the Used by Enders card in the right-hand rail — an Ender shown with a green (active) dot whose Set topics action includes the topic (see Which Enders use a topic).

Then let the data correct you. After the catalog is live, the share of conversations landing under No topic is the fastest signal that it does not match your real traffic — see When no topic is assigned. Tune individual topics on the workbench rather than by adding more topics; see Tuning a Gen AI topic on the workbench.

Creating a topic

  1. Click + Create in the top right of the Topics page.

  2. Pick the Type — Use Gen AI or Human Only for new topics. (Use Rules is not offered when creating a topic.)

  3. Enter a Name. The name is required.

  4. Fill in the type-specific field:

    • Use Gen AI — write a precise multiline Description of what the topic means. The description is required for this type, and both the topic name and description are sent to the AI when classifying a conversation, so write both with care. Example — Name: ATM, Description: Customer asks or tells information regarding ATM locations, issues, etc.

    • Human Only — no further fields. The topic is created with no automatic matching.

  5. Click Create to save.

    • For Human Only topics, the topic immediately becomes available in filters, dashboards, and reviewer pickers.

    • For Use Gen AI topics, the page automatically opens the topic on the workbench. The platform queues an initial structured definition and a first batch of historical samples to label in the background — both run automatically when the Gen AI topic is created.

Editing a topic

Click the pencil icon on any active topic to open it on the topic page. You can change the name, the categorization type, and the type-specific configuration (description for Gen AI, rules for an existing rule-based topic). Save with Update.

  • Use Rules appears in the type picker only when the topic being edited was saved as rule-based. Switching a rule-based topic to Use Gen AI or Human Only is supported, but switching another topic back to Use Rules is not offered.

  • Switching a topic to Human Only stops future automatic matching. If you switch from Use Rules, existing auto-assignments for that topic are cleared immediately while manually verified assignments are preserved. If you switch from Use Gen AI, existing assignments are not cleared automatically and remain on past conversations — remove them manually from each conversation if needed.

If you change a Gen AI topic's description while there are still Pending reviews queued for tuning, the platform asks you to confirm before saving: the change discards the pending review batch and prepares a new one based on the updated description. Confirm only if you want to start the review queue over.

Priority position and reordering

Each active topic has a Position value, shown as a numeric chip on the left of the row. Position ranks topics that have already been assigned to the same conversation: the lowest-numbered one becomes that conversation's "main" topic, which is what analytics views showing a single topic per conversation report. It does not affect which topics get assigned in the first place.

To reorder topics:

  • Drag-and-drop — works only when the table is sorted by Position and no quick filters are applied. The drag handle is hidden otherwise; hover the position cell to see the reason ("You must sort topics by position…" or "You can't use drag&drop ordering if filters are used…").

  • Set position manually — click the position chip on a row and enter a new value in the Update topic position dialog. Valid range is 1 to the total number of active topics.

Position cannot be edited from the Archived view.

Labels

The Labels column lets you tag topics for your own organization — for example, by initiative, owner, or theme. Click + Add labels to add existing labels or create a new one inline. Labels are independent of how a topic actually matches; they only help you find and filter topics in this list.

Searching and filtering the topic list

The quick-filter bar above the table narrows the list by:

  • Search — matches against topic names.

  • Labels — show topics carrying selected labels.

  • Type icons (robot for Gen AI, person for Human, list for Rules) — toggle each type on or off. All three are selected by default; turning any off counts as an active filter and disables drag-and-drop reordering.

Filters affect the visible list only and do not change topic configuration.

A Gen AI topic with pending review samples shows a clipboard icon in the row's action column. Click it to jump straight into the topic's launched review mode for that topic.

Accuracy column

The list shows an Accuracy column per topic. For a Gen AI topic, it shows the latest eval accuracy of the topic's currently active definition (formatted as a percentage). Rule-based and Human Only topics show an empty value because accuracy applies only to Gen AI definitions. Use this column to scan which Gen AI topics need attention on the workbench.

An Improvement ready chip appears next to the accuracy value when a non-active version of the topic has been evaluated and scored at least 5 percentage points higher than the active one. Click the chip to open the topic page, where the candidate's row in the Versions rail shows the accuracy gain and a Use this version button. The chip never appears on archived topics, on Rule-based or Human Only topics, or when the active version has no eval result to compare against.

Which Enders use a topic

Topics can be selected inside Enders (automations) — for example to categorize a conversation or as a trigger filter that decides which conversations an Ender runs on. The Topics feature surfaces those dependencies in two places, so you can see the impact before you change or archive a topic.

Used by Enders column (topics list). The list can show a Used by Enders column with the number of distinct Enders that depend on each topic. Each Ender is counted once, even when it references the topic in more than one place. An Ender counts when its topic selection includes the topic — whether you picked the topic directly, or through a bulk scope such as All Gen AI topics, All rule-based topics, or All topics — and also when the topic is referenced in the Ender's trigger filters. Enders that are inactive or archived still count, so the number reflects every Ender configured to act on the topic. Archived topics always show as unused.

Bulk scopes expand only to Gen AI and rule-based topics: All Gen AI topics to Gen AI topics, All rule-based topics to rule-based topics, and All topics to both. No bulk scope expands to Human Only topics. A Human Only topic therefore counts as used only when an Ender selects it directly or references it in a trigger filter — an Ender configured with All topics alone is not reported against a Human Only topic. (Selecting a Human Only topic directly does let an Ender apply it automatically; the bulk scopes simply never reach it.)

Used by Enders card (topic page). Open any saved topic and a Used by Enders card appears in the right-hand rail, listing each dependent Ender by name. A colored dot marks whether each Ender is currently active (green) or inactive (dimmed). When no Ender depends on the topic, the card reads Not used.

  • Ender names are clickable links that open the Ender for editing in a new tab only for users who also have the automation-management permission. Users without it see the Ender names as plain text.

  • When the topic is unused, users with automation-management also see a Create a new Ender link (hidden on archived topics).

  • Archived topics always show Not used.

When usage refreshes. Usage is recalculated when you open the Topics section and after you create, edit, import, archive, restore, or delete a topic. Enders open in a new tab, so editing one does not refresh an already-open Topics page — reload the page to pick up the change. (Navigating within the section, such as opening a topic and returning to the list, does not force a fresh usage load.) If a refresh fails, the column shows a dash (—) and the card shows Usage is currently unavailable instead of a possibly incorrect zero — reload the page to try again.

Version column

The Version column shows the version number of a Gen AI topic's currently active definition — for example 2 for the second definition in its history. It reflects the topic's own definition version and never the version of any Ender. Rule-based topics, Human Only topics, and Gen AI topics that do not yet have an active definition show a dash (—). This column is hidden by default; turn it on from the column control described below.

Choosing which columns appear

The topics list has a Configure columns button (the sliders icon) in the top-right toolbar, next to the Active / Archived toggle. Use it to tailor which optional columns are shown and in what order.

  • Name and Actions are always shown and cannot be hidden or reordered. Position is also fixed, but only in the Active view — the Archived view has no position column.

  • Type, Accuracy, Used by Enders, Version, Created, Updated, and Labels are configurable. Created and Version are hidden by default.

  • Tick or untick a column to show or hide it. At least one configurable column must stay visible — the last remaining one cannot be unticked.

  • Drag the handle next to a column name to reorder it.

  • Reset to default restores the default visibility while keeping your custom column order.

Your column choices are saved in the browser and shared between the Active and Archived views. Sort order is tracked separately: Active defaults to Position (ascending) and Archived to Name (ascending), and each view remembers its own sort. Switching views restores that view's sort and returns to the first page. If you hide the column a view is currently sorted by, the sort falls back to Position for Active or Name for Archived.

Archiving and restoring

Use the top-right Active / Archived toggle to switch views.

  • Archive a topic by clicking the archive icon on its row in the Active view, then confirming. Archived topics stop being applied to new conversations and are hidden from regular use, but historical assignments are preserved.

  • Restore an archived topic by clicking the restore icon on its row in the Archived view, then confirming. Restored topics resume matching.

The Archived view is read-only apart from the restore and delete actions — inline editing, reordering, and label changes are disabled.

Deleting a topic

Permanently delete a topic using the delete action on the row. A topic must be archived first — attempting to delete an active topic returns an error asking you to archive it. Deletion removes the topic and its configuration; existing topic assignments on past conversations are no longer linked to a topic name.

If you only want to stop using a topic without losing its history, archive it instead of deleting.

Importing and exporting

The top bar of the Topics page offers two file-based actions.

  • + Import — opens the Import topics dialog. Pick an Excel (.xlsx) file and click Upload. The platform creates the topics found in the file. If any topic name in the file already exists in the workspace, the import is rejected and no topics are created — the dialog shows which names conflicted. The check only looks at topics already in the catalog: duplicate names within the uploaded file are not detected, so if the file itself lists a name twice you get two topics. Imported Gen AI topics automatically queue an initial definition and a first review batch in the background, the same as topics created from the + Create dialog.

  • Export (Excel icon) — downloads the current topic catalog as an Excel file. Only Gen AI topics are included in the export, because rule definitions are not part of the supported import format.

Use import/export to set up topics in bulk, copy topics between environments, or back up the catalog.

Tuning a Gen AI topic on the workbench

Open a Use Gen AI topic for editing to open the single-page workbench. The topic configuration — Type, Name, and the multiline Description — is pinned at the top of the page. Everything that helps you tune the topic appears below the form on the same page, with no tab switching. The workbench refreshes itself: while background work is in flight (a definition is being generated, a review batch is being prepared, an eval is running, or an AI instruction is being created), the page silently polls and pulls in new versions, accuracy values, and pending-review counts the moment they land.

The form's Cancel and Update buttons stay disabled until you make an unsaved edit, so you can read the configuration without worrying about losing changes.

When a Gen AI topic is brand new and has not been saved yet, a banner reminds you to save first. The workbench panels described below appear once the topic exists.

Improve accuracy journey

Above the workbench cards, an Improve accuracy (by following these steps) stepper shows where the topic stands on a three-step path:

  1. Create topic — done as soon as the topic has any version.

  2. Review conversations — done after you have labeled any review sample.

  3. Get improvements — done after an improved candidate version exists.

If pending review samples are waiting (or review mode is otherwise ready to open), the Review conversations step doubles as a shortcut into review mode — click it to open the review dialog directly.

Once an improved version has been promoted and the topic is at rest, the stepper is replaced by the Best version yet panel described below, and the overline above the panel shortens to Improve accuracy.

Finishing a review

Labeling samples does not, on its own, change the topic's accuracy or produce a new version. Saving a label only updates the calibration dataset. Turning those reviews into a measured result takes a deliberate action, and in the review flow that action is the Finish review button at the bottom of the review dialog. (It is not the only way to measure a topic — the Improve and Evaluate buttons under Advanced do the same work one version at a time — but it is the one that finishes a review round.)

Finish review appears only when all of the following are true:

  • no pending samples are left in the queue,

  • at least 5 matching (Yes) and 5 non-matching (No) conversations have been confirmed, so the topic can be measured against balanced ground truth,

  • the topic has an active definition whose AI instruction has finished generating,

  • nothing else is currently being prepared on this topic (no definition generation, no review batch), and no evaluation is already running on the active definition. An evaluation you launched by hand on an older, non-active version does not block the button.

Until those conditions are met the button is not shown. The review dialog explains what is still missing in a notice above the samples — for example Add matching conversations to measure accuracy when one side is short, or Preparing review samples while a batch is still being built.

When the queue is empty and everything is in place, Finish review appears in the bottom-right corner of the dialog; its arrival is the signal that the round can be closed. Click it — you are asked to save or discard an unsaved reviewer note first — and the platform queues one round of work made up of two background jobs:

  • Measure the active definition. The live version is re-measured against the test subset of your confirmed dataset, using the Instruction the AI uses that production actually runs. This is the job that refreshes the headline accuracy and clears the Outdated marker.

  • Attempt one improvement. The active definition is re-run over the validation subset of the dataset and compared against your labels. The mistakes it makes there are analyzed and rewritten into a new candidate version, saved inactive. The same job then measures that candidate against the same test subset, so the candidate lands in the Versions rail with an accuracy figure directly comparable to the active one. This step is an attempt, not a promise — the conditions below describe when it produces nothing.

The two jobs are queued together and run independently, so their results appear on the workbench as each one finishes rather than in a guaranteed order — a candidate's accuracy can show up before the active version's refreshed number does. Let both status strips clear before you compare the two figures.

The dataset is split into those two subsets for you, and the split is uneven: roughly 20% of samples go to the test subset and the remaining 80% to validation. Improvement learns from the validation subset, and accuracy is measured on the test subset. Each sample is assigned by a stable hash of its conversation, whether it arrived through automatic sampling or you added it by hand. Overriding a high-confidence prediction is the one thing that overrides the hash — see Reviewing samples for exactly when.

The split is not a strict holdout on small datasets. A measurement wants at least 20 usable samples with at least 5 of each label. If the test subset alone cannot supply that, the platform tops it up with validation samples until it can. On a small dataset this means a candidate is scored partly on the very mistakes it was rewritten from, which flatters its accuracy. Because test is only a fifth of the dataset, this top-up lasts longer than it looks: at 40 confirmed samples the test subset holds about 8, so most of the 20 scored samples are borrowed from the material improvement just learned from. Expect the two subsets to separate cleanly only around 100 confirmed samples, reasonably balanced. Below that, treat a candidate's accuracy as a promising signal rather than an independent verdict, and keep labeling: the split becomes meaningful on its own as the dataset grows.

The improvement job is conditional, and the round ends quietly rather than failing when there is nothing to do:

  • Improvement is not queued at all when the dataset holds no human-labeled Yes or No sample in the validation subset — which a small hand-picked set can occasionally miss, since the subsets are decided by hash. You still get the refreshed accuracy; adding a few more examples clears it. Unclear labels are never used for improvement.

  • No candidate is created when the active definition already agrees with every one of your validation labels. There are no mistakes to learn from, so the refreshed accuracy is the only result.

  • The candidate measurement runs only if a candidate was actually created.

A round runs exactly once and never re-enqueues itself. You can run another one on the same dataset whenever you like — once the jobs settle, Finish review is offered again even if you have labeled nothing new — but rerunning against the same examples usually reaches the same conclusion, so label more samples first if you want the improvement step to have fresh material.

A Review completion queued toast confirms the request, the review dialog closes, and you return to the workbench, where the status alerts described below track the work. If the request cannot be accepted at all, the message is "Review could not be finished. Check the topic setup and try again."

If the topic becomes busy between opening the dialog and clicking the button — for example a review batch starts in the background — the action is refused with "Review cannot be finished yet. Complete any pending samples and wait for topic processing to finish." Wait for the workbench to settle and click Finish review again.

Promotion stays manual. Finishing a review only creates and measures the candidate. You still activate it yourself from the Versions rail (see Promoting a candidate version).

Description edits do not invalidate prior reviews. Reviewer yes/no labels judge the topic itself (anchored by its name), so refining the description to help the AI read the topic better is normally a refinement of the same intent and the platform continues to learn from the existing reviews the next time you finish a review.

Status strip on the active version

While the work started by Finish review — or any improvement or evaluation you launched by hand — is in flight, the active-version view shows a live status strip under the accuracy KPI:

  • **Improving from your N reviews…** with a progress bar — shown while the candidate definition is being generated. The count reflects how many reviewed samples the platform is currently learning from.

  • Measuring accuracy… with a progress bar — shown while an evaluation run is in progress.

  • Improvement failed with a Retry button — shown when an improvement run errored. Retry reruns the improvement on the active definition. This failure strip only appears when reviewed samples exist; a failed first-time definition still shows the Generate initial definition retry described in Background preparation status instead.

Background preparation status

The workbench shows background-preparation status inline rather than in a separate dashboard:

  • When no version exists yet, an in-progress message and progress bar appear in the accuracy card while the structured definition is being generated. If the definition fails, the message becomes a red alert with a Generate initial definition button to retry.

  • When a review batch is queued or running, a Preparing review samples alert appears under the accuracy KPI with a labeled-of-total progress count and a progress bar.

  • If a preparation step has been queued or running for more than ten minutes without progress, a Retry preparation button appears inside that alert. Use it sparingly — the original task may still finish later and produce duplicate work.

  • Next to the accuracy KPI on the active version, a Continue to review button appears as soon as there are pending samples to label, opening review mode for the topic. When no pending samples are waiting but the topic is ready (including the low-data case where no eligible conversations were found yet), the same slot shows a Create review batch button instead. It opens the same review dialog so you can broaden the calibration filter and recover.

Making sure a topic can be measured

Before the platform can measure a Gen AI topic's accuracy, it needs a balanced set of confirmed examples: at least 5 matching conversations (the topic applies) and at least 5 non-matching conversations (it does not). Automatic sampling usually finds both while it prepares the first review batch, but for rare topics it may not turn up enough matching conversations on its own — which is exactly when a topic could otherwise be measured against a lopsided, mostly non-matching set and report a misleadingly high accuracy.

When one side is short, the workbench and the review queue show Add matching conversations to measure accuracy (or Add non-matching conversations to measure accuracy) in place of the accuracy actions, with the explanation that the topic keeps categorizing conversations normally but cannot be measured yet. Underneath, a counter shows where both sides stand — for example "Matching: 2/5 · Non-matching: 7/5".

An Add matching conversations (or Add non-matching conversations) button opens the Conversations list in a selection mode scoped to the topic's calibration filter. This button appears only for users who can access Conversations. In that mode:

  • A bar at the bottom reads Finding examples for: {topic} and tracks progress — for example "2 of 5 confirmed · 3 selected".

  • Tick the conversations that clearly match (or clearly do not match) the topic.

  • Click Confirm N as matches (or Confirm N as non-matches) to save them as human-labeled examples.

  • Once the required count is reached, the platform returns you to the topic's review queue automatically. Below that count you stay in the conversations list so you can keep picking; use Cancel in the same bar to leave selection mode.

The same two buttons also sit at the top of the review queue whenever it holds samples and neither side is short, so you can top up either side at any point without waiting for a shortage message.

Selecting examples this way is not limited to the automatic sampling scope — you can pick any conversation your role can open, which lets you add clear-cut examples the sampler missed. Conversations without transcript text cannot be used as examples. An example added this way already carries the label you confirmed, and it is exempt from the explanation requirement described in Reviewing samples — you can change its label in the queue without writing a note, and add one later if you want to record your reasoning.

If your Ender uses Single-topic Categorisation. Review samples are drawn from the predictions your Enders have already produced on real conversations. Single-topic Categorisation records only the topic it matched, so a topic's pool of samples can consist of matching conversations alone — in which case a review batch comes back with matching samples only, and the non-matching side stays at zero no matter how many batches you create. Default Categorisation (multi-topic) also records the closest non-matches, so its batches usually arrive with both sides already present.

When that happens, use Add non-matching conversations as described above to add the missing side by hand. Once five non-matching examples are confirmed, measurement and improvement work exactly as they do for any other topic.

Once both sides reach the minimum, nothing runs by itself: the shortage notice disappears, Finish review appears in the bottom-right corner of the review dialog, and you start the measurement from there (see Finishing a review).

Accuracy KPI and version history

Below the configuration card, the workbench shows a headline Accuracy, formatted as a percentage, or Not evaluated yet when no measurement exists. It belongs to the version currently selected in the Versions rail, which is the active definition when you open the workbench but changes as soon as you click another row. The topics list Accuracy column always shows the active definition, so the two match only while the active row is selected — check which version is selected before comparing the workbench number to the catalog.

An Outdated marker appears next to that number when the calibration dataset has changed since the measurement was taken — which is what happens when you save or add a Yes or No sample. Only those two labels count: Unclear samples are left out of measurement entirely, so saving one does not age the accuracy. (Changing an existing Yes or No to Unclear does, because it removes a sample the number was built on.) Hover the marker for the hint "Accuracy may be outdated. Open Advanced, then click Evaluate to recalculate it." You do not have to go that far: the button next to the figure runs the same measurement on whichever version is selected, reading Recalculate accuracy when that version already has a number and Evaluate when it does not.

Finishing a review also refreshes the number, but only for the active version — that is the only definition Finish review measures. In the normal flow the active row is the one selected, so the marker clears on its own. If you have clicked an older row in the rail, that version keeps its own figure and its own Outdated marker no matter how many review rounds you finish; refresh it with the button beside the KPI, or leave it and select the active row again.

The right side of the workbench holds the Versions rail, listing every version of the topic's structured definition that the platform has produced or that you have edited manually, newest first. Each row shows:

  • The version number and an Active chip on the live version.

  • The creation timestamp and a By {author} line that names the user who saved a manual edit, followed by a provenance note for the two version kinds that need one: Original description and Manual. Versions produced by the platform's own generation and improvement runs carry no extra note. Older Gen AI topics that ran for a while before receiving their first structured definition get a system-authored Original description version. It preserves the description used before structured definitions as its AI instruction and lets you return to description-only categorization if the generated instruction performs worse. Topics created directly in the structured workflow do not get this extra version.

  • The latest eval result on that version — with its own Outdated marker when it predates the current calibration dataset — or Not evaluated, or a pending state while an eval is running.

  • An Instruction pending chip while the platform is still compiling the AI instruction for that version.

Hover or keyboard-focus a non-active version to reveal two actions: Activate swaps the active definition (and surfaces a Version activated toast on success), and the delete action removes a historical version that is not currently active or running an eval.

While the topic's preparation is busy — a definition is being generated, or a review batch is queued or running, including a retryable batch that has not been retried yet — both actions are disabled and explain why on hover ("Available once topic preparation finishes"). Retry preparation stays available the entire time so you can unstick a stalled batch without waiting for it. Activate is also disabled on a version whose AI instruction is still being generated.

Promoting a candidate version

When a non-active version has been evaluated and beats the active version by at least 5 percentage points in accuracy, its row in the Versions rail can gain two things the other rows do not have: the accuracy gain written out in green as active% → candidate%, and a green Use this version button in place of the usual Activate action. That green delta is what to look for when scanning the rail — there is no separate banner or label on the row.

The offer is made on one row at a time. If several versions clear the five-point bar, only the highest-scoring one gets the green delta and Use this version; the rest keep the ordinary Activate action even though they also beat the active version. So the absence of a green delta on a row does not mean that version is worse than the live one — read each row's own accuracy figure to compare them, and use Activate if you want a version other than the one the platform is pointing at.

The offer is a nudge, not an automatic action — review the candidate before promoting it. Activation always stays manual.

When Use this version is unavailable (for example because a definition is being generated or a review batch is queued or running), hover the button to see the reason — the hint stays visible even while the button itself is disabled.

Best version yet

Once an improved version has been promoted as the active definition and the topic is at rest, the journey stepper is replaced by a green Best version yet panel above the version list. The panel appears only when all of these are true:

  • the active version was created by an improvement run (not the initial definition or a regeneration),

  • the active definition has a measured accuracy,

  • nothing is currently queued, running, pending, or in a failed state,

  • no other candidate version beats the active one by enough to surface a Use this version offer.

The panel shows the current accuracy alongside one of two messages, depending on where the topic stands:

  • Below roughly 95% accuracy — **Accuracy improved to *X%*** with the hint "To push it higher, review more conversations — more labeled examples give the AI more to learn from." A primary Review more conversations button opens review mode so you can label a fresh batch.

  • At roughly 95% or above — **Accuracy is already excellent · *X%*** with the hint "Above ~95%, extra reviewing rarely moves the number — review more only if you spot mistakes." The same Review more conversations button is shown in outlined style so it does not look like a required next step.

A short note next to the button — "or fine-tune the definition under Advanced" — points to the Advanced card below if you would rather edit the structured definition than label more samples. When you create a new review batch (or otherwise put a job back in flight), the panel disappears and the journey stepper comes back.

Definition details and the slide-over drawer

The left side of the workbench shows the currently selected version. The summary card displays:

  • The latest eval result for the selected version, with the Accuracy, Samples, eval mode (Instruction the AI uses or Structured definition), and the TP / TN / FP / FN breakdown (true/false positives and negatives). The breakdown chips are clickable — see Drilling into evaluation results below.

  • A second card for the Last structured-definition eval when both eval modes have run on the same version.

  • The Instruction the AI uses — the text that is actually sent to the AI when classifying conversations.

  • The Change summary — your own short note about what changed in this version.

The full structured definition is not shown inline. Click View / edit full definition to open it in a slide-over Full definition drawer that overlays the page from the right. The drawer shows:

  • The Definition rule.

  • The Include when and Exclude when conditions.

  • The Boundary cases, each with a Scenario, a Yes / No label, and a Reason.

  • The Change summary.

  • A collapsible Saved topic description panel at the bottom that shows the customer-facing topic description that was in effect when the version was created. Use it to compare the tuned definition back to what the topic was originally meant to capture.

From the drawer, Edit definition opens an inline editor where you can change the Definition rule, the Include when and Exclude when conditions, the Boundary cases, and the Change summary. By default the AI instruction is regenerated automatically from your edits — toggle Edit AI instruction to override it manually instead. Save as new version creates a new manual version without activating it and closes the drawer; Cancel discards your edits and returns to the read view inside the drawer (the drawer stays open). Use the close (X) icon in the drawer header to close the drawer.

Drilling into evaluation results

The TP / TN / FP / FN chips on the Last eval result card are not just counts. Click a chip to open a dialog listing the individual conversations behind that number, so you can sanity-check what the AI actually got right and wrong before promoting a version.

Chip

Dialog title

What it lists

TP

Correct matches (TP)

Conversations expected to match the topic that the AI also matched.

TN

Correct non-matches (TN)

Conversations expected not to match that the AI also did not match.

FP

Incorrect matches (FP)

Conversations expected not to match that the AI matched anyway.

FN

Missed matches (FN)

Conversations expected to match that the AI missed.

Each entry shows the Conversation ID, the Expected label from your reviewed examples (Matches topic, Does not match topic, or Unclear), and the AI result in the same wording. Correct results are marked with a green edge and disagreements with an amber edge, so a long list stays quick to scan.

For the two disagreement outcomes — FP and FN — each entry additionally shows the model's Prediction confidence as a percentage and its Rationale, which is where you find out why the AI decided the way it did. Entries with no recorded rationale read No rationale provided. Correct results (TP and TN) show the labels only.

Opening a conversation. The Conversation ID is a link that opens the full conversation in a new browser tab, but only for users who can access Conversations. Users without that access still see the drill-down and everything in it — expected label, AI result, confidence, and rationale — with the ID shown as plain text instead of a link. The same rule applies to the conversation ID shown above the transcript in review mode.

When a chip is not clickable.

  • A chip whose count is 0 is never clickable, because there is nothing to list.

  • FP and FN list every recorded disagreement. If an evaluation stored no per-conversation results at all, its chips stay as plain counts with no drill-down.

  • TP and TN open only when the platform can prove the list is exact. If some conversations in the evaluated batch have no recorded result, the correct-result list cannot be reconstructed reliably, so the chip is shown with an information icon and reads "Some conversations have no evaluation result, so the exact conversation list is unavailable." on hover. The count itself is still accurate — only the conversation-by-conversation breakdown is withheld rather than shown incomplete.

Hover a clickable chip to confirm what it opens: Review correctly classified conversations for TP and TN, or Review disagreements with expected labels for FP and FN.

The dialog closes on its own once the result it was showing is no longer current — when you select a different version, or when a newer eval result lands for the topic. Reopen the chip on the refreshed card to see the updated list.

Drill-downs are offered on the Last eval result card only. The Last structured-definition eval card shows its accuracy and sample count without a per-conversation breakdown.

Improve and Evaluate

On the happy path, Finish review measures the active definition and attempts a candidate for you (see Finishing a review above). The Advanced card still exposes two manual buttons as a fallback — use them when you need to rerun an action on a specific historical version, or to launch an improvement or evaluation without going through review mode:

  • Improve queues an automated improvement run based on the currently selected version (the one shown in this details card, which may or may not be the active definition) using the reviewed samples. Improvement requires that the topic already has reviewed validation samples — if there are none, the button is disabled. It is also disabled while the AI instruction is still being generated.

  • Evaluate runs a test eval on the selected version against the reviewed test samples. By default the eval is run against the production Instruction the AI uses, which reflects real production behavior. Open the dropdown on the button and pick Evaluate structured definition to instead test the full structured definition — a diagnostic that tells you how much the compiled instruction diverges from the definition it was built from. Read it as a comparison, not a ceiling: the structured definition can score lower than the compiled instruction as well as higher. Confirm the prompt to start; the row in the version list shows a pending state until the eval finishes.

If a version was saved without an AI instruction (for example after a manual edit that left the field blank), the details card shows an information banner reading "The AI instruction is being generated. Activation, evaluation, and review sampling are blocked until it is ready." with a Generate AI instruction button. When a manual save triggers the platform to regenerate the instruction automatically, the workbench polls quietly in the background and refreshes itself once it is ready — no manual reload needed.

Review mode

Review mode opens from the workbench using the Create review batch button on the active version (when no pending samples exist) or the Continue to review button (when pending samples are waiting). You can also click the Review conversations step in the Improve accuracy journey when it is highlighted as clickable.

Review mode opens as a popup dialog over the workbench, with the topic name and a Close (✕) button in the header. The workbench stays mounted behind the dialog and refreshes itself when you close the dialog, so new versions, accuracy values, and pending-review counts produced during the session are reflected immediately.

Review mode stays reachable whenever the topic has a ready definition and nothing is running, including the low-data case where the last batch finished without finding any eligible conversations. In that case, opening review mode is how you broaden the calibration filter and recover.

The saved dataset is always reopenable

Review mode is not a one-way queue that empties and disappears. It shows the topic's whole calibration dataset — the samples still waiting for a decision and every sample anyone has already confirmed — so you can come back at any time to re-read a call, change your mind about a label, or write down why you decided what you decided.

A line at the top of the dialog states the scope and the split: "Review and edit the calibration samples for this topic." followed by **Pending reviews: N and Reviewed samples: *M***.

Move through the dataset with the Previous and Next buttons on either side of a **Sample X of *Y*** counter, where Y counts everything in the dataset, not just what is pending. Saving a label automatically advances you to the next sample, so labeling a fresh batch is still a straight run from the first sample to the last; on the final sample you stay put. Navigation is blocked while a sample is saving or its transcript is still loading, and if you have typed a reviewer note without saving it, stepping away asks Discard unsaved note? — "This note has not been saved. Choose a label to save it, or discard the changes to continue." The same prompt guards closing the dialog and clicking Finish review.

Because every sample stays reachable, revisiting a confirmed one shows the label you gave it and the note you wrote, both editable. Saving overwrites that sample in place — the new label, note, and reviewer replace the old ones, and the previous decision is not kept anywhere you can look it up. Re-confirming an already-reviewed sample without changing anything is a no-op and simply moves you on.

Samples cannot be deleted from the dataset one by one, and no action taken on the topic itself removes a reviewed sample — every removal from the workbench affects pending samples only. Three actions discard the pending batch: changing the Calibration filter in review mode, saving a changed Description on the topic, and changing the topic's categorization type. The first two warn you first when there is a pending batch to lose; a type change does not, so check the Pending reviews count before switching a Gen AI topic to another type.

That guarantee covers what you do inside the topic, not what happens to the conversations themselves. A sample points at a specific conversation, so if that conversation is deleted or removed by your retention policy, its sample goes with it — label, note and all — and the dataset shrinks. Expect the accuracy to move after a large deletion or a retention pass, and re-measure if the number matters.

Saving a Yes or No updates the calibration dataset, which is what makes the recorded accuracy Outdated until you measure again — see Accuracy KPI and version history.

Review progress journey

Once the current batch has any samples, the same Create topic → Review conversations → Get improvements stepper appears at the top of the review dialog. The connector between Review conversations and Get improvements is a live progress bar that fills as you label samples ("N of M reviewed"), so each decision pushes the topic visibly closer to the improvement step.

Calibration filter

A foldable Calibration filter panel sits at the top of the dialog body. Expand it to see the current scope and a short hint explaining when to narrow it (the topic only applies to part of your traffic — a specific queue, language, team, or date range) and when to leave it on All conversations (the topic applies broadly — narrowing it then yields fewer, less representative samples and can skew measured accuracy).

Edit the filter and click Apply filter & generate new samples to save the change. The filter is persisted to the topic, any previous pending review samples are dropped, the low-data state (if the last batch ended without eligible conversations) is cleared, and a fresh review batch is queued automatically under the new scope. If pending samples were about to be discarded, you'll be asked to confirm first. The button is disabled when the filter has not changed or when a definition or batch is currently running.

Reviewing samples

For each sample the dialog shows:

  • A single-line header that reads **Relevant transcript segment · Conversation ID N. For users who can access Conversations**, the conversation ID is a link that opens the full conversation in a new tab; for everyone else it is plain text. Either way the transcript below is fully readable, so reviewing is never blocked by this.

  • The full call transcript, with the segment that drove the model's decision visually highlighted and auto-centered on screen so you can read the context above and below it without scrolling first.

  • The model's Prediction confidence as a percentage and its Rationale — a short explanation of why the model picked Yes or No.

  • A Reviewer note (optional) field, pre-filled with whatever note the sample already carries. Its hint reads "Choose a label to save this sample." — the note is stored together with the label, so typing a note and navigating away without picking a label saves nothing.

  • Three decision buttons: Yes, No, and Unclear. The button matching the model's prediction is highlighted.

Use keyboard shortcuts to move quickly: 1 for Yes, 2 for No, 3 for Unclear, and Space to accept the suggested label as-is. Hotkeys are ignored while a typing field is focused or a dialog is open.

If you override a high-confidence prediction (the model was at least 75% confident and you pick the opposite label), the platform switches to a dedicated panel asking Why does this label apply? and will not save until you answer. Confidence overrides on Unclear, reviews that match the model's prediction, and samples you added by hand from the conversations list never require an explanation — a hand-picked example already carries a label you chose deliberately, so you can relabel it or leave its note empty freely.

Notes are worth writing even where they are optional. They are stored with the sample and travel with it into the improvement step, where the mistakes carrying a note are the ones the platform reaches for first when it rewrites the definition. Not every note gets that far, though: improvement reads only the Yes and No samples in the validation subset, so a note on an Unclear sample, or on one the hash put in the test subset, is kept with the sample but never shown to the AI. You cannot see which subset a sample is in while reviewing — write the note whenever you have something to say, and treat it as a record for your team first and an input to the AI second.

Answering the disagreement panel is the one thing that can move a sample into the validation subset, and it does so only when the label is genuinely new: the first time you review the sample, or when you change a label you had saved before. Re-confirming the same label on an already-reviewed sample, or going back only to reword its note, leaves the sample in whichever subset it was already in — so if the hash had put it in the test subset, the explanation you just wrote stays there and improvement never sees it. In practice this means a first pass through a batch is what feeds improvement; a later editing pass over samples you have already settled mostly does not.

If the current batch has too few likely matches to reach the five-per-side minimum, a warning appears above the samples: "We found only a small number of likely matching conversations. Review them first; afterward, you may need to add more to measure accuracy." When it is the non-matching side that is short, the same warning reads "We found only a small number of likely non-matching conversations…" instead. Either way, label the batch you have, then top up the side the warning names — Add matching conversations for the first wording, Add non-matching conversations for the second. See Making sure a topic can be measured above.

Reaching the last pending sample does not change the dialog into a summary screen. Because the whole saved dataset stays loaded, the sample browser remains on screen with your reviewed samples in it, and what the topic needs next is signalled around it rather than in place of it:

  • The counters at the top left — Pending reviews and Reviewed samples — are how you tell that the batch is done: pending reaches zero.

  • Add matching conversations and Add non-matching conversations sit in the top-right corner of the dialog whenever the dataset holds samples, not in a panel that appears at the end. Both need Conversations access.

  • Add conversations automatically, which asks the platform for another batch, joins them in that corner only when a batch can actually be built: nothing left pending, no preparation or measurement running, an active definition in place, and eligible conversations found by the previous batch. When any of that is missing the button is hidden, not greyed out, so there is no disabled control to inspect — if you cannot find it, one of those conditions is unmet. It is notably absent in the moment just before you finish a round: once both sides reach five confirmed examples and the topic has not been measured yet, the button disappears and Finish review takes its place. Finish the round, then ask for more conversations.

  • If the button is missing because the last batch found nothing eligible, edit the Calibration filter above and click Apply filter & generate new samples to broaden the scope; a fresh batch is queued as part of that action.

  • If matching or non-matching examples are still short of five once nothing is pending, a warning appears above the sample browser with the heading Add matching conversations to measure accuracy (or the non-matching wording), the shortage explanation, and the **Matching: n/5 · Non-matching: n/5 counter. Use the Add** buttons in the corner above it to top the short side up.

  • When the topic is ready to be measured, Finish review appears in the bottom-right corner of the dialog. That button is the signal that the round can be closed — see Finishing a review.

Finish review does not disappear once a round has finished, so you can start another round on the same dataset at any time.

The large Review complete, Evaluation in progress and No conversations are waiting for review. messages, along with Done for now, belong to the empty-dataset state and take over the dialog only when the topic has no saved samples at all — a topic you have not started reviewing yet, or one whose pending batch was discarded before anything was confirmed.

If the platform decides a batch needs to retry, the same Retry preparation action shown in the topic preparation panel is available from review mode as well.

Turning reviews into a new version

There is no separate "improve" button in review mode. The Finish review button at the bottom right of the dialog is the one action there that measures the topic and attempts a candidate version — it always refreshes the accuracy, but a candidate appears only when there is something to learn from, as Finishing a review explains in full. Everything you do in review mode before that — labeling, relabeling, noting, adding conversations — only shapes the dataset that Finish review then acts on.

The manual Improve and Evaluate buttons under Advanced remain available as a fallback when you need to rerun an action against one specific historical version.

Language

The platform writes generated topic definitions and AI instructions in the workspace's interface language, configured in System. Change the system language before generating a new definition if you want the AI-produced text in a different language; existing versions are not retranslated automatically.

Important notes

  • The Use Gen AI type only appears when a generative-AI provider is configured for the workspace. Without one, Human Only is the default type for new topics.

  • Whether a conversation can carry several Gen AI topics or only one is set per Ender, in the Prompt to use field of the Set topics action: Default Categorisation (multi-topic) assigns every justified topic, Single-topic Categorisation assigns at most one of the topics that action selected. The single-topic cap is per action, not per conversation: another Ender selecting different topics, or a rule-based or Human Only assignment, can still add a second topic. The field appears once the action selects any Gen AI topic, nothing is preselected, and the Ender cannot be saved until you choose.

  • A conversation can finish with no topic at all — always for misdials and accidental contacts, and under Single-topic Categorisation whenever nothing matches. This is a normal state, not an error; those conversations appear under No topic.

  • The topic catalog is flat — topics cannot be nested and have no parent/child relationship. Use Labels to group topics; a hierarchy from a legacy manual catalog has to be flattened.

  • Position is a reporting rank, not a matching rule. It is never sent to the AI and does not change which topics are assigned — it only decides which of a conversation's assigned topics is treated as its main one in analytics.

  • Only + Import rejects duplicate topic names, and only against topics already in the catalog (case-insensitive, archived topics included) — duplicate names inside the uploaded file itself are not detected. Creating topics one at a time applies no name check, and no path detects topics that overlap in meaning. Review the catalog yourself: merge duplicates, or separate them with Exclude when and Boundary cases.

  • There is no cap on catalog size. An Ender using a bulk scope such as All Gen AI topics evaluates every active topic of that type against every conversation it runs on, so archive topics you no longer use rather than leaving a long tail of rare ones active. An Ender that selects topics individually only evaluates those.

  • Archiving keeps a topic's historical assignments and restores them intact, but an archived topic is skipped by topic analytics — conversations that carried only that topic are reported under No topic until it is restored.

  • A newly created topic is not classified against anything until an Ender's Set topics action includes it. Gen AI and rule-based topics can be included directly or through a bulk scope; a Human Only topic counts only when it is selected individually, as bulk scopes never expand to it. Saving rules on a rule-based topic rematches the last 30 days once but does not tag new conversations on its own.

  • Used by Enders is reliable only when it reads Not used. A non-zero count also covers trigger-filter references and Enders that are inactive or archived, so confirm a topic is live by opening it and checking that an active (green-dot) Ender's Set topics action includes it.

  • Review samples come from predictions your Enders already produced. Under Single-topic Categorisation only matching topics are recorded, so review batches can arrive with matching samples only; add the other side with Add non-matching conversations to reach the five-per-side minimum.

  • Use Rules is a legacy type. It is hidden from the type picker for new topics and is only offered as an option when editing a topic that was already saved as rule-based.

  • Position values affect prioritization across analytics views, not just this list — change them with care.

  • Drag-and-drop reordering requires the table to be sorted by Position with no filters applied and the Active view selected.

  • Archiving is reversible; deletion is not, and requires the topic to be archived first.

  • Creating or importing a Gen AI topic automatically queues background preparation work for the initial definition and the first review batch.

  • The workbench refreshes itself while a definition, review batch, eval, or AI instruction is being prepared — there is no manual refresh button.

  • Activate and the delete action on a version are disabled while a definition is being generated or a review batch is queued or running (including a retryable batch). Retry preparation stays available throughout so you can unstick a stalled batch.

  • Changing the description of a Gen AI topic that still has pending review samples discards that batch on save and prepares a new one.

  • Changing the Calibration filter in review mode and clicking Apply filter & generate new samples persists the filter, drops the current pending batch, and queues a fresh batch against the new scope in one step. If pending samples would be discarded, you are asked to confirm first.

  • Review mode always shows the topic's whole calibration dataset — pending samples and every confirmed one. Page through it with Previous / Next and the **Sample X of Y counter to relabel a sample or edit its Reviewer note (optional) at any time. A note is only stored when you also pick a label; navigating away with an unsaved note prompts Discard unsaved note?**.

  • Individual samples cannot be deleted, and relabeling one overwrites it in place — the previous label and note are not kept. Three actions drop pending samples and keep every reviewed one: changing the Calibration filter, saving a changed Description, and changing the topic's categorization type. Only the first two warn you first.

  • The Improve action needs reviewed validation samples to work; review the full pending batch before triggering an improvement run.

  • Measuring a Gen AI topic's accuracy requires at least 5 matching and 5 non-matching confirmed conversations. When one side is short, the workbench shows Add matching conversations to measure accuracy (or the non-matching wording) and an Add matching conversations button that opens the Conversations list so you can confirm the missing examples. The add-examples buttons appear only for users with Conversations access.

  • Nothing measures or improves the topic on its own. Finish review — shown once the queue is empty, both sides have five confirmed examples, an active definition with a ready AI instruction exists, no definition or review batch is being prepared, and no evaluation is running on the active definition — queues two independent background jobs: one measures the active definition, the other attempts an improvement and then measures the candidate it produced. They finish in no guaranteed order, so let both settle before comparing the numbers. A hand-launched evaluation on an older version does not hold the button back.

  • The improvement job is not queued when the dataset holds no human-labeled Yes/No sample in the validation subset, and produces no version when the active definition already agrees with every validation label. In both cases the refreshed accuracy is still the outcome. A round never re-enqueues itself, but you can click Finish review again once it settles — even without new labels, though a rerun on the same examples usually lands in the same place.

  • The round only creates and measures a candidate — promotion always stays manual through Use this version on the row showing the green active% → candidate% gain, or the per-row Activate action. When several versions beat the active one, only the best of them carries that green offer; the others are promoted with the ordinary Activate.

  • The dataset splits roughly 20% test / 80% validation. Improvement learns from validation and accuracy is measured on test, but the split is not a strict holdout while the dataset is small: a measurement pulls in validation samples whenever the test subset cannot supply 20 samples with 5 of each label. Because test is only a fifth of the dataset, expect that top-up until around 100 confirmed samples; until then a candidate's accuracy is partly scored on what it learned from and reads high.

  • While the round is running, the active-version view adds a status strip (**Improving from your N reviews…, Measuring accuracy…, or Improvement failed with Retry**) below the accuracy KPI — the accuracy row and the review-dataset button stay where they are.

  • Saving or adding a Yes or No calibration sample marks the recorded accuracy Outdated — the previous figure stays visible so you never lose the last known number, and the marker clears when the next measurement finishes. Unclear samples are not measured, so saving one leaves the accuracy alone.

  • The Best version yet panel only appears when the active version came from an improvement run and no work is pending or queued — putting a new job back in flight returns the workbench to the journey stepper.

  • Retry preparation is only meant for batches that have stalled for more than ten minutes — the original task may still finish, so use it sparingly.

  • Generated definitions and AI instructions are written in the workspace language set under System at the time of generation. Switch the system language first if you need them in a different language.

  • The Accuracy column on the topic list always reports the active definition's full-history latest eval result. The workbench Accuracy KPI reports the version selected in the Versions rail, which starts on the active one — so the two agree until you click a different version. Both stay empty for Rule-based and Human Only topics, and for Gen AI topics whose active definition has never been evaluated.

  • The green accuracy gain and Use this version button on a candidate row, and the Improvement ready chip on the topic list, are nudges — not automatic actions. Activation always stays manual.

  • The TP / TN / FP / FN chips on the Last eval result card open a drill-down of the conversations behind each count. FP and FN entries also show the model's confidence and rationale; TP and TN entries show the expected label and AI result only.

  • TP and TN drill-downs open only when the exact conversation list can be proven complete. When it cannot, the chip shows an information icon explaining that the list is unavailable — the count stays accurate. Chips with a count of 0 never open.

  • Conversation IDs in the drill-down and in the review-mode transcript header are clickable links only for users with Conversations access. Without it, the labels, confidence, and rationale are still fully visible — only the link is unavailable.

  • Drill-downs are available on the Last eval result card only, not on the Last structured-definition eval card. The dialog closes automatically when you switch versions or when a newer eval result arrives.

  • The Used by Enders column and the topic-page card count every Ender whose topic selection includes the topic — directly or through a bulk scope such as All Gen AI topics — plus Enders that reference it in a trigger filter. Bulk scopes cover only automatically matched types, so Human Only topics count only when selected directly or referenced in a trigger filter — never via All topics. Inactive and archived Enders still count, each Ender is counted once, and archived topics always show as unused. Links to open or create an Ender appear only for users with automation-management permission.

  • If a usage refresh fails, the list shows — and the topic page shows Usage is currently unavailable rather than a false zero. Reload the page to retry.

  • The Version column shows the active Gen AI definition's version number, never an Ender's version. It is empty (—) for Rule-based, Human Only, and not-yet-defined Gen AI topics, and is hidden by default.

  • Use Configure columns to show, hide, and reorder optional columns. Name and Actions stay fixed (with Position fixed in the Active view only), Created and Version are hidden by default, at least one configurable column must stay visible, and the preferences are shared across the Active and Archived views.

Did this answer your question?