Skip to content

References: Prompt Testing and Research Dialogue

  1. Software testing - Wikipedia - Offers an accessible overview of Software testing, including definitions, methods, examples, limitations, and related concepts. This foundation helps students reason carefully about testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  2. Interview (research) - Wikipedia - Offers an accessible overview of Interview (research), including definitions, methods, examples, limitations, and related concepts. Its examples help students evaluate evidence for testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  3. Leading question - Wikipedia - Offers an accessible overview of Leading question, including definitions, methods, examples, limitations, and related concepts. It supplies useful context for decisions about testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  4. Interviewing Users (2nd ed.) - Steve Portigal - Rosenfeld Media - Shows how to plan interviews, build rapport, ask neutral questions, probe stories, capture evidence, and synthesize responsibly. Its sustained treatment supports work on testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  5. Software Testing: A Craftsman's Approach (5th ed.) - Paul C. Jorgensen - Auerbach Publications - Provides systematic test design, boundary analysis, model-based testing, coverage concepts, failure diagnosis, and repeatable quality practices. Its cases illuminate tradeoffs involved in testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  6. Working with Evals - OpenAI - Introduces test data, evaluation criteria, graders, repeated runs, and comparison workflows for measuring model behavior instead of relying on impressions. Its methods give teams a starting point for testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  7. Demystifying Evals for AI Agents - Anthropic - Explains agent evaluation design, realistic tasks, outcome and process graders, repeated trials, transcript review, and analysis of variable behavior. Its comparisons clarify choices involved in testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  8. Writing Survey Questions - Pew Research Center - Explains questionnaire development, pretesting, wording, response formats, order effects, and common sources of measurement error. Its framework strengthens responsible work on testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  9. User Interviews: How, When, and Why to Conduct Them - Nielsen Norman Group - Covers interview planning, neutral facilitation, follow-up questions, limitations of self-report, and appropriate uses for attitudinal research. Its examples show how evidence informs testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.

  10. Prompt Engineering Guide - OpenAI - Presents practical techniques for writing clear instructions, supplying relevant context, using examples, structuring tasks, and improving model reliability through iteration. Its implementation advice helps teams practice testing prompt behavior and conducting neutral, probing, adversarial, and reflective dialogue.