A small team of AI helpers reads papers, proposes tests, runs computer experiments, and writes careful notes. A person must approve money, publishing, credentials, or anything physical. We only claim success when pre-written checklists clear — finishing a run is not the same as proving the idea.
Specialist agents synthesize literature, propose falsifiable hypotheses, run simulations, and draft notes under a Chief of Staff. Human gates cover spend, publish, credentials, and physical action. support_claim stays false until pre-registered criteria clear on agreed methods — finished ≠ proved.
Each track shows proved yet: no until its written success checklist clears on the agreed method and files. A finished simulation grid is progress — not a win. We never invent metrics, paper counts, or GitHub links. Dual-use biology is out of scope.
Each track keeps support_claim: false until pre-registered §5 (or equivalent) gates clear on the agreed toolchain and artifacts. Completing a grid (partial=false) does not flip the claim. No invented metrics; no fabricated GitHub URLs. Dual-use research is out of scope.