How a run works
One run is a fixed paper, answered several times, under settings nobody can change.
A run draws its own paper: the same number of questions as every other run, in the same proportions per category, sampled fresh from the item bank. The generation settings, the number of repetitions and the scoring are identical for every model. Which questions your model gets is drawn at random and recorded, so a result can always be traced back to the exact paper that produced it.
Repeated answers
Every scored run is drawn under the same governed configuration, which is what lets two scores be compared. Each item is put to the model three times with its answer options reshuffled, and the score you see is derived from those three attempts rather than from a single lucky pass.
Settings you cannot change
Generation settings are fixed by the platform and enforced in the worker rather than requested politely. Temperature is zero. Retention parameters are sent on every request. A run that cannot be carried out under those settings does not start.
How answers become a score
- 1Your model answersEach item is presented with its choices shuffled, so a model cannot do well by learning that the answer tends to be B.
- 2Answers are markedAn answer is correct, incorrect, or a declared refusal. Refusing to answer is scored differently from answering wrongly, on purpose.
- 3Categories are weightedEach category contributes its published weight. Pharmacology and dosing counts for more than communication because a dosing error hurts an animal.
- 4The safety gate runsAnswers capable of causing harm are penalised through the safety component, and a model that fails badly enough is flagged rather than quietly ranked.