How do you isolate evaluation trials so earlier runs cannot inflate scores?

Instruction: Identify state leakage and show how each trial starts from a controlled environment.

Context: Make evaluation isolation concrete for files, databases, caches, and external tool behavior.

Updated

Official answer available

Read the opening below, then unlock the full answer and practical guidance.

I’d give every trial a fresh copy of its starting state and a unique namespace for files, database records, and tool requests...

Your preparation path

Work through these questions in order. Read the answer aloud, then explain it in your own words.

Related Questions