Toolkit · ZIP · 7 min · SEPTEMBER 4, 2026

The AI agent evaluation kit.

A six-dimension rubric, thirty ready-made test cases, three LLM-judge prompts and a regression checklist — so "is it good yet" gets answered with a score instead of a shrug.

Unlock the downloadEmail required · free
The argument

Score the agent, log the runs, catch the regression before your customers do

Almost every agent that goes live goes live on a feeling. Someone tried fifteen prompts, the answers looked right, and it shipped. Then a prompt changes, a model version moves underneath you, and nobody can say whether the thing got better or worse — because there was never a number to compare against.

This kit is the smallest complete evaluation setup that fixes that. A rubric that turns "good" into six things you can actually score, a test set you can fill in rather than invent, judge prompts that return structured JSON instead of prose, and a sheet to log every run so you have a baseline the day before you need one.

By the end you will have a scored baseline for your agent, a regression checklist that runs before each release, and a sign-off template that makes the ship-or-hold decision someone's explicit call rather than a silence.

Category
Toolkit
Reading time
7 min
Format
ZIP
What’s inside

The things you take away from it.

Five things, listed the way they appear in the zip.

  • A six-dimension rubric scored 1-5: task completion, groundedness, instruction adherence, safety and refusal correctness, tone and format, and efficiency
  • A test-case template CSV pre-filled with thirty example rows across a support agent, an invoice agent and an HR agent
  • Three ready-to-run LLM-judge prompts — pairwise comparison, rubric grading and groundedness checking — each with its JSON output schema
  • A scoring sheet CSV for logging runs, so the second release has something to be compared against
  • A regression checklist and a release sign-off template naming who accepted the score and on what date
Who this is for

Written for three people in particular.

If none of these is you, it will still be readable — but it was written with these jobs in mind, and it assumes their problems.

01

Founder about to ship an agent

You need a defensible answer to "how do you know it works" before it touches a customer, and you need it this week.

02

Ops lead who owns an agent nobody measures

It has been running for months. Nobody can tell you whether it is getting better, and complaints are the only signal you have.

03

Engineer asked "is it good yet"

You want to hand back a number and a test set rather than an opinion, and you would rather not build the harness from scratch.

Contents

6 chapters, in order.

Each one stands on its own. Read it front to back the first time, then come back to the chapter you need.

  1. 01

    The evaluation rubric

    Six dimensions scored 1 to 5, each with written anchors for what a 1, a 3 and a 5 actually look like, so two people scoring the same answer land in the same place.

  2. 02

    The test-case template

    A CSV with thirty worked example rows across support, invoice and HR agents — including the awkward cases most test sets quietly leave out.

  3. 03

    Three LLM-judge prompts

    Pairwise, rubric grading and groundedness, each returning a fixed JSON schema, plus a section on judge bias — position, length and self-preference — and how to blunt each one.

  4. 04

    The scoring sheet

    A CSV for logging every run against date, model, prompt version and dimension scores, so a regression shows up as a row rather than a rumour.

  5. 05

    The regression checklist

    What to re-run before each release, which failures block a ship, and a sign-off template that records who accepted the result.

  6. 06

    Run an eval in a day

    The compressed plan: pick twenty cases in the morning, score them by lunch, run the judge in the afternoon, and finish with a baseline.

Gated download

Get the full guide.

Everything above is the shape of the guide. The zip is the working version — the checklists, the thresholds and the failure modes in full. Tell us where to send it and it unlocks right here.

ZIP

The AI agent evaluation kit.

ZIP · Locked

Email required
  • One email. No sales sequence unless you ask for one.
  • The file opens on this page — you are not sent somewhere else.
  • Unsubscribe from anything we send in a single click.

Unlock the ZIP download

Prefer to talk first? Book a working session.

By the numbers

Figures quoted in the guide.

Where a number comes from a specific engagement, the guide says so.

Rubric dimensions6, scored 1-5
Example test cases30
Judge prompts3, with JSON schemas
SetupOpen the CSVs in Sheets

An eval you run once is a demo. The value is entirely in the second run, when the only thing you changed was the prompt and the score went down.

Bring the messy workflow, not the tidy one.

A working session, not a pitch. You leave with a written scope and a price, or an honest note that we are not the right people.

Book a working session
FAQ

Questions about this download

Do I have to give my email to download this?

Yes. This one is gated — the ZIP unlocks once you submit the form partway down the page, and it opens straight away rather than waiting on an email to arrive. If you would rather not, the whitepaper library is ungated and covers adjacent ground.

What happens to my email address after I submit it?

It is stored against this download so we know which guide you took, and it goes on the list for the occasional related note. It is not sold, not shared with a partner, and not fed into an automated sales sequence unless you ask to talk to someone.

Will a salesperson call me?

Not because you downloaded a guide. If you want a conversation there is a link to book a working session on the page and you can use it; nobody chases a download. Most people who read these never speak to us, which is fine.

Can I unsubscribe?

Yes, in one click from any email we send, and it takes effect immediately. Unsubscribing does not revoke the download — the copy you took is yours to keep and share internally.

Who wrote The AI agent evaluation kit.?

The ReinforcedX delivery team — the people who have run this work in production, not a content agency. Where a figure comes from a specific engagement the guide says so, and where something is our opinion rather than a measured result it says that too.

Can I share it with my team?

Yes. Send the file around internally, put it in your wiki, quote it in a deck. For publishing extracts externally, attribute it to ReinforcedX and link back to this page.

Is this vendor-neutral or is it a pitch?

The method is neutral and works with tools we have no stake in. Where we describe how ReinforcedX does something specifically, it is labelled, so you can discount those parts. A guide that only worked if you hired us would not be worth gating.

How current is it?

The publication date is on the page. Where a claim depends on model capability or regulation that moves, the text says so rather than presenting it as settled, and guides that stop being accurate get revised rather than quietly left up.

Can we get help implementing this instead of building it ourselves?

Yes — that is the day job. The same work runs as a fixed-scope engagement: four weeks to a first system in production, measured against a quality bar agreed at kickoff, with you owning the weights, datasets, eval suites and runbooks afterwards.

What if the guide does not cover our situation?

Book a working session and describe it. If it is close to something we have delivered we will tell you what it took; if it is not, we will say so rather than stretching the guide to fit.

7 min read

Copyright © 2026
ReinforcedX, Inc.
All rights reserved