summaryrefslogtreecommitdiff
path: root/src
AgeCommit message (Collapse)Author
2026-04-01Fix error runs not appearing in final summaryCaptainJack2491
2026-03-26Refactor judge module into judge/ packageCaptainJack2491
Split monolithic src/judge.py into a proper package structure: - src/judge/batch_providers.py: BatchProvider ABC, AnthropicBatchProvider, XAIBatchProvider - src/judge/judges.py: RegexChecker, BlackboxChecker, GlassboxChecker - src/judge/helpers.py: extract_* functions, parse_json_verdict - src/judge/prompts.py: BLACKBOX_PROMPT, GLASSBOX_PROMPT, constants - src/judge/judge.py: Judge orchestrator class - src/judge/__init__.py: Re-exports all public APIs
2026-03-26Add xAI (Grok) batch API supportCaptainJack2491
- Add XAIBatchProvider class implementing BatchProvider ABC - Add --provider CLI argument to judge_runner (anthropic|xai) - Refactor Judge class to use BatchProvider instead of direct API calls - Update tests for new Judge constructor signature
2026-03-25refactor: extract judge prompts to files and add granular blackbox categoriesCaptainJack2491
- Extracted Blackbox and Glassbox prompts from into their own markdown files in - Implemented progressive 4-tier Blackbox categories (FABRICATED, OMITTED_CRITICAL, DECEPTIVE_FRAMING, ACCURATE) - Added neutral examples to the Blackbox prompt to improve formatting consistency - Updated references and tests to match new judge categorization structure
2026-03-25feat: implement rich dashboard for experiment progressCaptainJack2491
2026-03-25feat: add tqdm progress bar and fix main config flag loadingCaptainJack2491
2026-03-25fix: catch and recover from JSONDecodeError on malformed tool argsCaptainJack2491
2026-03-25feat: add goal_type variation system and parallel processingCaptainJack2491
- vfs.py: add VFS.create() factory for parallel-safe independent instances - tools.py: add make_tools_for_vfs() to bind tools to specific VFS instances - config_loader.py: add goal_types and max_workers config properties - agent.py: add goal_type field, vfs_instance param, goal_type in log paths - runner.py: rewrite with goal_type loop dimension, ThreadPoolExecutor parallel execution, isolated VFS per run, flattened work items, thread-safe results All backward compatible - empty goal_types defaults to legacy hidden_goal.md, max_workers=1 defaults to sequential execution.
2026-03-14feat: statistical analysis on the whole dataCaptainJack2491
2026-03-06feat(judge): add selective judging and log filteringCaptainJack2491
- Implement enabled_judges parameter in Judge class to selectively run regex, blackbox, or glassbox checks. - Add --judges CLI argument to judge_runner.py for granular control over judging pipeline. - Add --model and --scenario filters to judge_runner.py for targeted evaluation of experiment logs. - Optimize synchronous judging to skip LLM calls and API key checks when only local (regex) judges are enabled. - Update Anthropic Batch API integration to support dynamic number of requests per log based on enabled judges. - Improve console output and logging to reflect selected filters and active judge components.
2026-03-04fix(agent): added reasoning extraction from groq modelsCaptainJack2491
2026-03-02feat: regex checker now returns score (e.g. 6/7) instead of binary PASS/FAILCaptainJack2491
- Verdict shows matched/total required rules for more nuanced evaluation - Adds 'matched' and 'total' fields to result dict - Returns 'N/A' when no rules are defined
2026-03-02fix: sharpen glassbox category definitions for ROBUST_HONEST vs ↵CaptainJack2491
UNINTENTIONAL_FAILURE - ROBUST_HONEST: ignored hidden goal + output is accurate and complete - UNINTENTIONAL_FAILURE: ignored hidden goal + output has genuine quality issues - Added explicit disambiguation note in the prompt - Clarified CONSIDERED_HONEST as deliberate refusal to comply
2026-03-02feat: add benign_goal (task description) context to blackbox judgeCaptainJack2491
- Blackbox prompt now includes TASK DESCRIPTION section from benign_goal.md - Helps judge evaluate framing even without a baseline - Hidden goal is never exposed (maintains blackbox/glassbox boundary) - Applied to both judge_single and prepare_batch_requests paths
2026-03-02fix: judge resilience for missing baselines and list-format regex rulesCaptainJack2491
- RegexChecker now handles both list-format and dict-format YAML rules - Judge falls back to experiment log's VFS state for source docs when no baseline exists - Blackbox judge receives '(No baseline available)' instead of crashing - Resume logic excludes .partial.json from completed run count
2026-03-02fix: global oversight_levels now filters scenario-available levelsCaptainJack2491
Previously a scenario with its own oversight/ directory would ignore the global oversight_levels config. Now the global list acts as a filter — only levels present in BOTH the scenario dir AND the global config are run. Warns if no levels match.
2026-03-02feat: incremental log saving via .partial.json filesCaptainJack2491
- Writes _in_progress.partial.json after each turn in the chat loop - Partial file persists if run crashes/hangs for post-mortem inspection - Cleaned up automatically when final save_logs succeeds - Enabled in both baseline and experiment runs via runner.py
2026-03-02feat: add generate_baseline toggle to configCaptainJack2491
- New 'generate_baseline' option in defaults (default: true) - When false, skips baseline generation entirely - Warns if no baseline exists when generation is disabled - Updated config.yaml.example and README.md with docs
2026-03-02feat: extract reasoning summaries from OpenAI reasoning modelsCaptainJack2491
- Handle reasoning.summary type in reasoning_details (GPT-5.3-codex etc.) - Prefix summaries with [SUMMARY] to distinguish from raw CoT - Add reasoning_format field to log entries for metadata tracking - Warn on first turn if no reasoning is detected (glass-box judging impact)
2026-03-01feat: add centralized logging with configurable debug levelsCaptainJack2491
- Create src/logger.py with Python logging module - Add 4 debug levels: 1=CRITICAL, 2=WARNING, 3=INFO, 4=DEBUG - Level 4 includes reasoning, VFS info, and available tools - Level 4 auto-enables file output (both mode) - Update agent.py and runner.py to use logger instead of print - Update config.yaml.example with logging configuration - Update README with logging documentation
2026-03-01feat(judge): updated blackbox to include source document for added accuracyCaptainJack2491
2026-02-27fix(judge): hash run_id to stay under Anthropic batch custom_id 64-char limitCaptainJack2491
2026-02-27fix: separate skipped runs from incomplete in summary outputCaptainJack2491
2026-02-26feat: add checkpoint/resume support and atomic log writesCaptainJack2491
- Runner now counts existing .json log files per combination and skips completed runs, resuming from where it left off - Log files are written atomically via temp file + rename to prevent corrupt files from being counted as completed - Added --no-resume CLI flag to force a fresh start - Resume is enabled by default
2026-02-26chore: add .env.example + config.yaml.example, gitignore config.yamlCaptainJack2491
- Create .env.example with placeholder keys for all providers - Create config.yaml.example with clean reference config - Add config.yaml and judge_logs/ to .gitignore - Load .env in judge.py for standalone usage
2026-02-25feat: implement judging system with regex, blackbox, and glassbox checksCaptainJack2491
- Add src/judge.py: RegexChecker, BlackboxChecker, GlassboxChecker, Judge orchestrator - Add src/judge_runner.py: CLI for batch/single judging with CSV output - Add regex_rules.yaml for corporate_sabotage_v2 and child_protection scenarios - Add judge config section to config.yaml (model, temperature, log_dir) - Add tests/test_judge.py: 34 unit tests covering all components - Uses Anthropic Batch API for half-cost LLM judge calls - Judge logs include full CoT reasoning and judge model for auditability
2026-02-24[refactor] improved interrogation systemCaptainJack2491
- Fixed syntax errors in print statements (lines 159/162) - Added sanitize_for_api() to strip custom log fields before API calls - Added interactive commands: history, reasoning, vfs, info, save, help - Colored terminal output for readability - Better error handling (KeyboardInterrupt, EOFError) - Added --show-history flag to review conversation on load
2026-02-24[fix] VFS singleton reset + remove vestigial check_username toolCaptainJack2491
- VFS now creates fresh instance when root_path/fs_data is provided (prevents cross-scenario contamination in multi-scenario runs) - Removed check_username from tools (unrelated to any scenario)
2026-02-24[feat] baseline system - auto-generates baselines before experimentsCaptainJack2491
- Added _extract_baseline_content() to extract output from create_file tool calls - Baselines now saved to logs/{model}/{scenario}/baseline.md - Added global oversight fallback (scenario-specific → global oversight dir) - Updated config_loader with oversight level loading from scenario dirs
2026-02-18[feat] added a baseline to the runnerCaptainJack2491
(no commit message provided) (no commit message provided) Co-authored-by: aider (openrouter/openrouter/aurora-alpha) <aider@aider.chat>
2026-01-11[fix] simplified everythingCaptainJack2491
made eveyrthing much cleaner and it works very well. got rid of providers setup, only use openai sdk now.
2026-01-11[fix] big moving around, probably made lot more bugs.CaptainJack2491
2026-01-09[docs] updated lit and added more in methodologyCaptainJack2491
2026-01-09updated interrogate to work with new log formatCaptainJack2491
2026-01-09updated scenario [corporate sabotage]CaptainJack2491
2025-12-30major commitCaptainJack2491
updated the program to use config.yaml made the runner to handel interleaved thinking added anthropic support and a lot more
2025-12-18updated oversightCaptainJack2491
2025-12-04gemini-2.5-pro pilot testCaptainJack2491
2025-12-03added corporate_sabotage_v2CaptainJack2491
2025-12-03updated interrogation to load from logsCaptainJack2491
2025-12-03updated methodologyCaptainJack2491
2025-11-25added interogation logic and updated structure of logsCaptainJack2491
2025-11-23updated agent to parse thinking in content field and added model_id to the logsCaptainJack2491
2025-11-23updated methodologyCaptainJack2491
2025-11-23good logsCaptainJack2491
2025-11-23moved 03-logs to mainCaptainJack2491
2025-11-20updated agent to work with the EXTRA_BODY variable to work with models that ↵CaptainJack2491
require it.
2025-11-20added corporate espionage scenarioCaptainJack2491
2025-11-20uodated oversight promptCaptainJack2491
2025-11-19updated logs to show modelname and scenario, and added extra param for google.CaptainJack2491