Prompt Quality Evaluator
Run the Prompt Quality Evaluator MicroSim Fullscreen
About This MicroSim
You cannot improve a response you cannot judge. This MicroSim trains judgement directly: five prompt and response pairs, four quality dimensions, and an expert rating waiting behind a button.
Each example is rigged to fail in exactly one way. One response wanders off topic. One is fluent and completely wrong. One stops halfway through the task. One says a correct thing at catastrophic length. The fifth is genuinely good, which is harder to spot than you would think.
The scoring is not the point. The gap between your number and the expert's number is the point.
How to Use
- Read the prompt and response at the top of the drawing area.
- Set all four sliders from 1 to 10. Commit to a number before you check — guessing after the fact teaches nothing.
- Press "Check My Ratings" to reveal the expert scores. Your rating shows as a blue circle, the expert's as a black triangle, and the bar between them is green if you were within 2 points, amber within 4, red beyond that.
- Read the one-line lesson under the score bars, then press "Next Prompt" for the next example.
Iframe Embed Code
You can add this MicroSim to any web page by adding this to your HTML:
1 2 3 4 | |
Lesson Plan
Learning Objective
Students will be able to assess the quality of AI prompts and responses using the four evaluation dimensions of relevance, accuracy, completeness, and conciseness.
Bloom's Level: Evaluate (L5) — assess, judge
Grade Level
High school through adult learners.
Duration
15-20 minutes
Prerequisites
Students should know the four quality dimensions by name from Chapter 2.
Activities
- Cold calibration (8 min): Students work through all five examples alone, committing to ratings before revealing expert scores. They record their match percentage for each.
- Compare the misses (7 min): In pairs, students find the example where they were furthest from the expert and argue for their own number. Sometimes the student is right — the goal is a defensible rating, not agreement.
- Apply to their own work (5 min): Students rate a response they received earlier in the week on all four dimensions and identify which single dimension to fix first.
Discussion Questions
- Example 2 is fluent, well-structured, and factually wrong. Which dimension catches that, and why is it the hardest one for a reader to check?
- Can a response score 10 on accuracy and 2 on completeness at the same time? What would that look like?
- Conciseness is not the same as brevity. What is the difference?
- Why might two expert raters disagree by 2 points and both be right?
Assessment
- Can the student rate a new prompt and response pair within 2 points of an expert on at least three of four dimensions?
- Can the student name which dimension a given failure belongs to?
- Can the student explain why "the response was bad" is not a usable diagnosis?