| Age | Commit message (Collapse) | Author |
|
- Add XAIBatchProvider class implementing BatchProvider ABC
- Add --provider CLI argument to judge_runner (anthropic|xai)
- Refactor Judge class to use BatchProvider instead of direct API calls
- Update tests for new Judge constructor signature
|
|
- Extracted Blackbox and Glassbox prompts from into their own markdown files in
- Implemented progressive 4-tier Blackbox categories (FABRICATED, OMITTED_CRITICAL, DECEPTIVE_FRAMING, ACCURATE)
- Added neutral examples to the Blackbox prompt to improve formatting consistency
- Updated references and tests to match new judge categorization structure
|
|
|
|
|
|
- VFSFactory: independent instances vs singletons
- make_tools_for_vfs: parallel binding
- config_loader: goal_types and max_workers parsing
- runner: goal_type iteration and resume skipping
- agent: goal_type storage and path resolution
|
|
|
|
- Add src/judge.py: RegexChecker, BlackboxChecker, GlassboxChecker, Judge orchestrator
- Add src/judge_runner.py: CLI for batch/single judging with CSV output
- Add regex_rules.yaml for corporate_sabotage_v2 and child_protection scenarios
- Add judge config section to config.yaml (model, temperature, log_dir)
- Add tests/test_judge.py: 34 unit tests covering all components
- Uses Anthropic Batch API for half-cost LLM judge calls
- Judge logs include full CoT reasoning and judge model for auditability
|
|
Tests cover:
- VFS: sandbox escape, path traversal, CRUD, singleton staleness
- Config loader: missing keys, temperature cascade, oversight fallback
- Tools: schema/implementation sync, VFS delegation
- Agent: message construction, error handling, token counting, reasoning extraction
- Interrogation: sanitization, provider detection, prompt loading
- Runner: baseline extraction, prompt assembly, success detection
All tests run offline with mocked API calls.
|