A web agent recognizes an evaluation question and finds its published answer key. How should you handle the result?
Instruction: Separate answer correctness from valid evidence of the capability the benchmark was intended to measure.
Context: Address evaluation awareness, runtime contamination, and the limits of blocking leaked benchmark material.
Updated
Official answer available
Read the opening below, then unlock the full answer and practical guidance.
I’d flag the run as contaminated under the benchmark’s stated rules. A correct answer obtained from an answer key does not demonstrate the intended research skill...
Your preparation path
Work through these questions in order. Read the answer aloud, then explain it in your own words.
1. Start with the foundations
Build the vocabulary and explain the core decisions.
2. Apply it to a real workflow
Practice diagnosis, validation, and everyday tradeoffs.
- What signals matter most when evaluating tool-using agents? Free sample
- Your tool-using agent passes answer-level evals and still makes bad tool calls. How would you fix the blind spot? Member answer
- Production feedback says answers feel less trustworthy even though offline metrics are flat. How would you investigate? Member answer
3. Prepare for senior discussions
Explain failure boundaries, recovery, and production choices.
- Human reviewers are expensive, but model graders miss subtle safety issues. How would you balance the two? Member answer
- Two coding agents have different scores and sandbox limits. How would you compare them fairly? Member answer
- An answer tells your LLM grader to award full marks. How would you keep the evaluation trustworthy? Member answer
Related Questions
-
easy
-
easy
-
easy
-
easy
-
easy
-
easy