Skip to content

References: Evaluation Rubrics, Scoring, and Evidence

  1. Rubric (academic) - Wikipedia - Offers an accessible overview of Rubric (academic), including definitions, methods, examples, limitations, and related concepts. This foundation helps students reason carefully about building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  2. Likert scale - Wikipedia - Offers an accessible overview of Likert scale, including definitions, methods, examples, limitations, and related concepts. Its examples help students evaluate evidence for building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  3. Weighted arithmetic mean - Wikipedia - Offers an accessible overview of Weighted arithmetic mean, including definitions, methods, examples, limitations, and related concepts. It supplies useful context for decisions about building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  4. Measuring the User Experience (3rd ed.) - Tom Tullis and Bill Albert - Morgan Kaufmann - Covers metric selection, rating scales, task measures, confidence intervals, benchmarking, study design, and evidence communication. Its sustained treatment supports work on building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  5. Measurement Theory and Applications for the Social Sciences - Deborah L. Bandalos - Guilford Press - Explains constructs, scales, reliability, validity, scoring, factor models, and responsible interpretation of social-science measurements. Its cases illuminate tradeoffs involved in building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  6. AI Risk Management Framework - NIST - Provides the Govern, Map, Measure, and Manage functions for addressing validity, reliability, transparency, privacy, fairness, accountability, and other AI risks. Its methods give teams a starting point for building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  7. Working with Evals - OpenAI - Introduces test data, evaluation criteria, graders, repeated runs, and comparison workflows for measuring model behavior instead of relying on impressions. Its comparisons clarify choices involved in building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  8. Demystifying Evals for AI Agents - Anthropic - Explains agent evaluation design, realistic tasks, outcome and process graders, repeated trials, transcript review, and analysis of variable behavior. Its framework strengthens responsible work on building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  9. User Research Methods - Usability.gov - Organizes practical methods for planning studies, learning about users, evaluating designs, analyzing evidence, and communicating research findings. Its examples show how evidence informs building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.

  10. Benchmarking - Google Analytics Help - Explains peer-group benchmark ranges, normalized and absolute metrics, comparison limits, and responsible interpretation of relative business performance. Its implementation advice helps teams practice building fair rubrics, anchored scales, weighted criteria, confidence ratings, and evidence records.