fix: change default judge_backend from openclaw to api - #357
Merged
Conversation
The openclaw backend spins up a full OpenClaw agent with tools for judging. The agent was going rogue - running exec commands and exploring the filesystem instead of just returning JSON scores. This caused 'Failed to parse judge JSON response' errors and 0% scores on most tasks. The api backend does direct API calls without agent tools, which is more reliable for the simple scoring task. Root cause: Judge transcripts showed the agent calling find, ls, etc. instead of outputting scores. 128 parse failures in a single run.
Contributor
Code Review SummaryStatus: No Issues Found | Recommendation: Merge Solid, targeted fix. The change correctly addresses the root cause — switching the default Files Reviewed (1 file)
Reviewed by claude-4.6-sonnet-20260217 · 73,409 tokens |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The
openclawbackend spins up a full OpenClaw agent with tools for judging. The agent was going rogue — runningexeccommands and exploring the filesystem instead of just returning JSON scores.Evidence from live benchmark (Kimi K2.6):
find,ls, exploring workspaceSolution
Change the default
judge_backendfromopenclawtoapi.The
apibackend does direct API calls without spawning an agent with tools. This is more reliable for the simple scoring task.Files Changed
scripts/lib_grading.py— Change default parameter ingrade_task()and_grade_llm_judge()Testing
Need to run a benchmark with this fix to verify judge responses are now parseable JSON.