Design assessment logic, adaptive algorithms, prompts and scoring in isolation. Run them, measure them, compare them — then promote only what works.
Question counts, types, adaptive rules, scoring methods and passing logic — all as experiment variables.
Generation, evaluation, feedback and improvement prompts are versioned with the assessment.
Every run is stored as an experiment so you can answer: is this design better than the last one?