Prompt Evaluation Methods: Testing Prompts Before They Ship
A prompt that works on one example might fail on ninety others. Here is how I evaluate prompts before they ship. A prompt that produces one good output might produce ninety bad ones on other inputs. The only way to know whether a prompt works is to evaluate it on a set of test inputs and score the outputs against criteria. Prompt evaluation is the discipline that separates prompts that ship from prompts that work only on the examples their author tried. After building evaluation sets for several LLM features, I have a method that catches bad prompts before they reach users. This guide covers the method. Building an Evaluation Set An evaluation set is a collection of inputs paired with the criteria the output should meet. I build sets of twenty to fifty inputs that span the range the prompt will see in production: easy cases, typical cases, edge cases, and adversarial cases. A set of only easy cases produces a prompt that fails on hard inputs. A set that includes the hard cases produces a prompt that survives them. Easy cases confirm the prompt handles the basic task. Typical cases confirm the prompt handles everyday inputs. Edge cases confirm the prompt handles unusual but valid inputs. Adversarial cases confirm the prompt handles hostile or tricky inputs. Defining Scoring Criteria For each input, I define what a correct output looks like, in terms that can be checked.