Core ideas

Experiments and candidates

Test better ways of running the agent.


Experiments compare candidate agent setups against the current setup.

excellent experiment start

Expected result: Excellent records the candidate diff and the tasks used for comparison.

Candidates may change instructions, model selection, tool descriptions, context, checking, retrying, and stopping rules.

If a candidate sees the held-out tasks or lacks required evidence, it is invalid.