---
name: design-a-model-test
description: Design a small comparative test for model choices on a concrete task, using consistent inputs, an explicit rubric and verifiable results. Use before selecting a model for a repeatable feature.
---

# Design a model test

This is a public exercise written for this page, not a claim of benchmarks already performed.

Define the actual task and what a usable result must do. Pick a small, representative set of inputs the user is permitted to use, including difficult and ambiguous cases. Use synthetic or redacted examples when real data is not appropriate, and state how that limits the test.

Set evaluation criteria before generating results: factual support, required format, completeness, appropriate uncertainty and failure behaviour, as relevant. Distinguish a hard failure from a stylistic preference. Use the same inputs and comparable settings across candidates. Record the exact model/version, date and settings if a test is run.

Evaluate outputs against the rubric. Where feasible, review without model names to reduce brand preference. Measure latency and provider-reported cost only when execution records exist; include retries. Do not infer price or speed from reputation, and do not invent measurements.

A request to design a test does not authorize paid API calls, uploads of private data or production changes. If execution is authorized, retain enough non-sensitive evidence to reproduce the comparison. State sample limits and avoid declaring a universal winner from a small task-specific test.

Return the test inputs or their specification, rubric, recording format and a decision rule. Recommend a candidate only after results support it. Otherwise leave the choice open and identify the next check.
