A repeatable evaluation framework that runs prompts, datasets, and scoring logic across model versions.