EvalStudio

An experimentation lab for AI-powered assessments.

Design assessment logic, adaptive algorithms, prompts and scoring in isolation. Run them, measure them, compare them — then promote only what works.

Everything configurable

Question counts, types, adaptive rules, scoring methods and passing logic — all as experiment variables.

Prompts as first-class

Generation, evaluation, feedback and improvement prompts are versioned with the assessment.

Built for comparison

Every run is stored as an experiment so you can answer: is this design better than the last one?