| Age | Commit message (Collapse) | Author | |
|---|---|---|---|
| 2026-04-10 | refactor(judge): simplify to single judge per prong, auto-create sync clients | CaptainJack2491 | |
| BREAKING CHANGE: Judge API now uses single blackbox_model/glassbox_model instead of lists with aggregation. Changes: - Judge auto-creates sync clients from model configs (supports anthropic, openai, xai) - Removed multi-model aggregation support (aggregate_results method removed) - New: create_sync_clients_for_models() and get_supported_providers() - Updated judge_runner.py to use new single-judge API - Added scripts/judge_comparison.py for testing judge model pairs - Updated v2_redesign notes to reflect implementation status | |||
| 2026-04-02 | docs: add v2 redesign notes with publication plan | CaptainJack2491 | |
| - 01: v1→v2 evolution history, results, limitations, rationalization hypothesis - 02: judge system state, single judge per prong, 20% validation protocol - 03: research plan analysis, Option B (two-phase) selected - 04: publication plan — paper structure, three-layer analysis strategy, Gemini ceiling effect handling, budget, venue targets, timeline Key decisions: - Two-phase design (540 runs) over full factorial (1,620) - Single paper: v1 exploratory → Study 1 (oversight) → Study 2 (framing) - Three-layer analysis: pooled → per-model → sensitivity excluding ceiling models | |||
