# Building scorecards

Writing items, levels and feedback instructions that grade consistently, and the two mistakes that make a scorecard useless.

Source: https://replicatelabs.ai/docs/admin/scorecard-config
Last updated: 2026-09-03

---

A scorecard is the standard a conversation is graded against. Getting it right is the highest-leverage configuration work on the account, because every roleplay result and every management read downstream inherits its quality.

## The structure

<figure class="dfig">
<div class="dframe">
<div class="danat">
<div class="danat-row"><b>Item</b><span>One capability. A scorecard is a set of them. Aim for the handful that actually decide outcomes in your business rather than a complete taxonomy of selling.</span></div>
<div class="danat-row"><b>Levels</b><span>Under each item, what each grade looks like as observable behaviour.</span></div>
<div class="danat-row"><b>Feedback instruction</b><span>What the rep is told at each level. This is where your organisation's coaching voice lives.</span></div>
</div>
</div>
<figcaption><b>Item, levels, feedback.</b> Most of the quality is in the levels.</figcaption>
</figure>

## Levels must be behavioural

This is the whole craft. A level written as an adjective cannot be graded consistently by anyone, human or otherwise.

<figure class="dshot">
<img src="/img/docs/admin-scorecards.webp" alt="Scorecard configuration showing capability items and their score ranges" width="1440" height="900" loading="lazy" />
<figcaption><b>Scorecards on an account.</b> Several short scorecards, each grading one kind of conversation, beat one long one that grades all of them badly.</figcaption>
</figure>


<div class="dsplit">
<div class="dpanel no">
<div class="dpanel-h">Not gradeable</div>
<p><b>Item:</b> Discovery quality<br/>
<b>Level 3:</b> Good discovery<br/>
<b>Level 2:</b> Adequate discovery<br/>
<b>Level 1:</b> Weak discovery</p>
</div>
<div class="dpanel yes">
<div class="dpanel-h">Gradeable</div>
<p><b>Item:</b> Quantified the problem<br/>
<b>Level 3:</b> Got a number from the buyer and confirmed how it was arrived at<br/>
<b>Level 2:</b> Got a number, did not test it<br/>
<b>Level 1:</b> Established a problem exists, no number attempted</p>
</div>
</div>

The test: could two different people, reading only the level descriptions, watch the same conversation and land on the same grade? If not, rewrite until they could.

## Feedback instructions

Write what you would say to the rep, in your organisation's voice, at that level. Not the definition of the level again.

- **At the top level**, name what they did so it is repeatable. "You asked for the figure and then asked how they got to it. That second question is what makes the number usable."
- **At the middle**, name the one thing missing. Exactly one.
- **At the bottom**, give the words. A rep at level one usually does not know what the behaviour sounds like, and telling them to do it more does not help.

## Two mistakes

<div class="dnote warn">
<span class="dnote-h">Too many items</span>
<p>A twenty item scorecard produces a grade nobody reads and feedback nobody acts on, and it makes every roleplay feel like an exam. Six to ten items covering the capabilities that actually decide your deals will change more behaviour than a complete one.</p>
</div>

<div class="dnote warn">
<span class="dnote-h">Editing it once results exist</span>
<p>Scores are only comparable while the scorecard is unchanged. Edit items or levels and every earlier grade was produced against a different standard, with no offset that repairs it. Tune the scorecard hard in the first few weeks, then freeze it, and record the date you froze it so any report spanning that line can say so.</p>
</div>

## Match the scorecard to the conversation

A scorecard grades one kind of conversation. Binding a discovery scorecard to a negotiation roleplay marks reps down for not doing something they were never attempting, which reads to them as the tool being broken and is a fast way to lose a rollout.

If your team runs four distinct conversation types, you need four scorecards, not one long one.

## Where to start

Do not write one from scratch. Take the qualification standard your organisation already uses, the one in your deal review template or your methodology material, and turn each requirement into an item. If a requirement cannot be turned into observable behaviour, it was never a standard, and finding that out is useful on its own.
