diff options
| author | CaptainJack2491 <jayrupnakawala@gmail.com> | 2026-04-10 14:50:40 +0100 |
|---|---|---|
| committer | CaptainJack2491 <jayrupnakawala@gmail.com> | 2026-04-10 14:50:40 +0100 |
| commit | 9155f664a2db0abe841d0d7b9b059ffbe2290c18 (patch) | |
| tree | 8f59187b2b7677727eb19d990b7de377d38e6bad /config_dryrun.yaml | |
| parent | a72e0d6e93e6ea8d44b4ca86d51903bfbf4e21d5 (diff) | |
refactor(judge): simplify to single judge per prong, remove multi-model aggregation
- Remove multi-model aggregation support from Judge class
- Single blackbox judge: claude-sonnet-4-20250514 (anthropic)
- Single glassbox judge: gpt-4.1 (openai)
- Remove aggregate_results method and all aggregation-related code
- Simplify batch collection to single result per run
- Update config_dryrun.yaml with new judge configuration
Diffstat (limited to 'config_dryrun.yaml')
| -rw-r--r-- | config_dryrun.yaml | 19 |
1 files changed, 16 insertions, 3 deletions
diff --git a/config_dryrun.yaml b/config_dryrun.yaml index 1bb99da..cd362ba 100644 --- a/config_dryrun.yaml +++ b/config_dryrun.yaml @@ -12,19 +12,19 @@ models: provider: openrouter extra_body: reasoning: - effort: "high" # Using low for dry run to save tokens + effort: "high" exclude: false - id: z-ai/glm-5 provider: openrouter extra_body: reasoning: - effort: "high" # Using low for dry run to save tokens + effort: "high" exclude: false - id: deepseek/deepseek-v3.2 provider: openrouter extra_body: reasoning: - effort: "high" # Using low for dry run to save tokens + effort: "high" exclude: false goal_types: @@ -51,3 +51,16 @@ logging: format: "[{level}] {message}" output: both file: logs/dry_run.log + +judge: + log_dir: logs/judge_dryrun + + blackbox: + model: claude-sonnet-4-20250514 + provider: anthropic + temperature: 0 + + glassbox: + model: gpt-4.1 + provider: openai + temperature: 0 |
