# Evaluation starter kit

No live model runs, real user workloads, invoice data, or model-quality conclusions are included. The three cases are invented test fixtures, labeled as such. Replace or supplement them with redacted representative cases from your own work before a production decision.

1. Keep `evaluation_cases.json` unchanged for a cohort; commit its SHA256 with your result records.
2. Copy `evaluation_template.json` to a new file. Fill provider/route/configuration fields before any separately authorized model calls. Both candidates start at explicit medium effort for one controlled condition; this does not mean equal compute. Run a separate tuned condition if needed.
3. Record every real attempt, including failed and uncertain attempts. Add rows for retries; never replace a failed attempt with a successful result.
4. Grade the declared criteria. Do not run arbitrary generated code outside a dedicated, authorized sandbox. This kit does not execute returned code.
5. Reconcile actual charges. Set charge_reconciled true only with a matching billing_evidence reference; missing values remain null. Record review time separately.
6. Run `python summarize_eval.py your_records.json`. The summary never selects a winner or validates that supplied evidence is genuine.

Start with fresh, separate histories for each candidate. Token budgets, cache settings, tool implementations, prompts, SDK/transport version and geography belong in the run note. A native API test does not certify a gateway route. Cross-model history reuse is a different experiment.

The empty template is intentionally executable as an input to the summarizer. Its output has zero observed attempts and unknown costs; these are absence-of-observation values, not model performance measurements.
