Running HLE LLM-as-judge evaluation...
Judging question with gpt-5.4 repeat 1...
Judging question with claude-opus-4-6 repeat 1...
Judging question with gemini-3.5-flash repeat 1...
Judging question with gpt-5.4 repeat 2...
Judging question with claude-opus-4-6 repeat 2...
Judging question with gemini-3.5-flash repeat 2...
Judging question with gpt-5.4 repeat 3...
Judging question with claude-opus-4-6 repeat 3...
Judging question with gemini-3.5-flash repeat 3...
Reward: 0.333333
Evaluation complete.
