Documentation · Replicate Labs

The product, page by page. Every page is also served as Markdown, so the same words reach the models a team works in.

Roleplay

Scorecards

How a conversation is graded: items, levels, and feedback instructions, and why a scorecard edit resets your history.

A scorecard is the standard a conversation is graded against. It is built per account from that organisation's own methodology and qualification bar, so a grade means something specific rather than "the model liked it".

#The structure

ItemOne capability being assessed. "Quantified the buyer's problem", "confirmed the decision process", "handled the price objection without discounting". A scorecard is a set of these.
LevelsUnder each item, what each grade actually looks like, described as observable behaviour rather than adjectives. This is what makes grading repeatable rather than a vibe.
Feedback instructionWhat the rep should be told when they land at a given level. This is why the feedback reads like your organisation's coaching rather than generic advice.
Item, levels, feedback. The levels are the part that has to be behavioural. "Good discovery" is not gradeable; "asked for a number and got one" is.

#Why it is not a generic sales rubric

The items reflect how your organisation sells. A team on a problem-centred methodology is graded on whether the problem was quantified and its impact confirmed. A team on a value-centred framework is graded on the business issue and who owns it. Same conversation, different scorecard, legitimately different grades.

That is the point. A scorecard that graded every methodology the same way would be measuring generic conversational competence, which nobody's number depends on.

#Comparability

A scorecard edit is a new baseline

Scores are only comparable while the scorecard behind them is unchanged. If items or levels are edited, earlier grades were produced against a different standard, and plotting them on one chart is misleading, however tempting the line looks.

There is no offset that repairs this. Treat an edit as starting a new baseline, and say so on any report that spans the change.

This matters most during the first few months, which is exactly when a scorecard is most likely to be tuned. If you are going to change it, change it early and then leave it alone.

#Who sees your scores

You, and the managers of the teams you are in. Managers can also read the roleplay conversation itself, not only the grade. See What you can see.

#Building or changing one

Scorecards are configured per account. See Building scorecards for what makes a good item and the traps to avoid.

Last updated 3 September 2026

View as Markdown