| Age | Commit message (Collapse) | Author | |
|---|---|---|---|
| 2026-03-06 | feat(judge): add selective judging and log filtering | CaptainJack2491 | |
| - Implement enabled_judges parameter in Judge class to selectively run regex, blackbox, or glassbox checks. - Add --judges CLI argument to judge_runner.py for granular control over judging pipeline. - Add --model and --scenario filters to judge_runner.py for targeted evaluation of experiment logs. - Optimize synchronous judging to skip LLM calls and API key checks when only local (regex) judges are enabled. - Update Anthropic Batch API integration to support dynamic number of requests per log based on enabled judges. - Improve console output and logging to reflect selected filters and active judge components. | |||
| 2026-02-25 | feat: implement judging system with regex, blackbox, and glassbox checks | CaptainJack2491 | |
| - Add src/judge.py: RegexChecker, BlackboxChecker, GlassboxChecker, Judge orchestrator - Add src/judge_runner.py: CLI for batch/single judging with CSV output - Add regex_rules.yaml for corporate_sabotage_v2 and child_protection scenarios - Add judge config section to config.yaml (model, temperature, log_dir) - Add tests/test_judge.py: 34 unit tests covering all components - Uses Anthropic Batch API for half-cost LLM judge calls - Judge logs include full CoT reasoning and judge model for auditability | |||
