# Scorecards

How a conversation is graded: items, levels, and feedback instructions, and why a scorecard edit resets your history.

Source: https://replicatelabs.ai/docs/roleplay/scorecards
Last updated: 2026-09-03

---

A scorecard is the standard a conversation is graded against. It is built per account from that organisation's own methodology and qualification bar, so a grade means something specific rather than "the model liked it".

## The structure

<figure class="dfig">
<div class="dframe">
<div class="danat">
<div class="danat-row"><b>Item</b><span>One capability being assessed. "Quantified the buyer's problem", "confirmed the decision process", "handled the price objection without discounting". A scorecard is a set of these.</span></div>
<div class="danat-row"><b>Levels</b><span>Under each item, what each grade actually looks like, described as observable behaviour rather than adjectives. This is what makes grading repeatable rather than a vibe.</span></div>
<div class="danat-row"><b>Feedback instruction</b><span>What the rep should be told when they land at a given level. This is why the feedback reads like your organisation's coaching rather than generic advice.</span></div>
</div>
</div>
<figcaption><b>Item, levels, feedback.</b> The levels are the part that has to be behavioural. "Good discovery" is not gradeable; "asked for a number and got one" is.</figcaption>
</figure>

## Why it is not a generic sales rubric

The items reflect how your organisation sells. A team on a problem-centred methodology is graded on whether the problem was quantified and its impact confirmed. A team on a value-centred framework is graded on the business issue and who owns it. Same conversation, different scorecard, legitimately different grades.

That is the point. A scorecard that graded every methodology the same way would be measuring generic conversational competence, which nobody's number depends on.

## Comparability

<div class="dnote warn">
<span class="dnote-h">A scorecard edit is a new baseline</span>
<p>Scores are only comparable while the scorecard behind them is unchanged. If items or levels are edited, earlier grades were produced against a different standard, and plotting them on one chart is misleading, however tempting the line looks.</p>
<p>There is no offset that repairs this. Treat an edit as starting a new baseline, and say so on any report that spans the change.</p>
</div>

This matters most during the first few months, which is exactly when a scorecard is most likely to be tuned. If you are going to change it, change it early and then leave it alone.

## Who sees your scores

You, and the managers of the teams you are in. Managers can also read the roleplay conversation itself, not only the grade. See [What you can see](/docs/managers/what-you-can-see).

## Building or changing one

Scorecards are configured per account. See [Building scorecards](/docs/admin/scorecard-config) for what makes a good item and the traps to avoid.
