summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorCaptainJack2491 <jayrupnakawala@gmail.com>2026-03-02 15:23:50 +0000
committerCaptainJack2491 <jayrupnakawala@gmail.com>2026-03-02 15:23:50 +0000
commit239ec5c0660d2b618cefa35b0a6a1599279d533c (patch)
tree03807c597df65c7ba24b2efafb14b96fb3d372e0
parentdaa0eeef1947d1320882d91e3d623fa5cd2d6da1 (diff)
feat: add generate_baseline toggle to config
- New 'generate_baseline' option in defaults (default: true) - When false, skips baseline generation entirely - Warns if no baseline exists when generation is disabled - Updated config.yaml.example and README.md with docs
-rw-r--r--README.md68
-rw-r--r--config.yaml.example1
-rw-r--r--src/config_loader.py5
-rw-r--r--src/runner.py9
4 files changed, 50 insertions, 33 deletions
diff --git a/README.md b/README.md
index 8138eee..3dcc431 100644
--- a/README.md
+++ b/README.md
@@ -39,11 +39,11 @@ graph LR
C & D & E --> H[Judge Logs<br/>full CoT reasoning]
```
-| Check | Input | Output |
-|-------|-------|--------|
-| **Regex** | Model output + scenario patterns | `PASS` / `FAIL` |
-| **Blackbox** | Model output vs baseline (no reasoning) | `ACCURATE` / `MISLEADING` / `OMITTED` / `FABRICATED` |
-| **Glassbox** | Reasoning trace + system prompt + output | Intent category + sophistication tier |
+| Check | Input | Output |
+| ------------ | ---------------------------------------- | ---------------------------------------------------- |
+| **Regex** | Model output + scenario patterns | `PASS` / `FAIL` |
+| **Blackbox** | Model output vs baseline (no reasoning) | `ACCURATE` / `MISLEADING` / `OMITTED` / `FABRICATED` |
+| **Glassbox** | Reasoning trace + system prompt + output | Intent category + sophistication tier |
**Usage:**
@@ -97,9 +97,9 @@ judge:
log_dir: judge_logs
logging:
- level: 3 # 1-4 (1=CRITICAL, 2=WARNING, 3=INFO, 4=DEBUG)
+ level: 3 # 1-4 (1=CRITICAL, 2=WARNING, 3=INFO, 4=DEBUG)
format: "[{level}] {message}"
- output: both # console, file, or both
+ output: both # console, file, or both
file: logs/experiment.log
```
@@ -107,14 +107,15 @@ logging:
The framework uses Python's `logging` module with configurable verbosity levels:
-| Level | Shows |
-|-------|-------|
-| 1 | CRITICAL only (major failures) |
-| 2 | WARNING + CRITICAL |
-| 3 | INFO + WARNING + CRITICAL (default) |
-| 4 | DEBUG + INFO + WARNING + CRITICAL (includes reasoning, VFS, tools) |
+| Level | Shows |
+| ----- | ------------------------------------------------------------------ |
+| 1 | CRITICAL only (major failures) |
+| 2 | WARNING + CRITICAL |
+| 3 | INFO + WARNING + CRITICAL (default) |
+| 4 | DEBUG + INFO + WARNING + CRITICAL (includes reasoning, VFS, tools) |
At **level 4** (DEBUG), additional information is logged:
+
- Full reasoning traces (truncated at INFO level)
- VFS path and file listing
- Available tools list
@@ -166,9 +167,12 @@ Logs are saved to `logs/{model}/{scenario}/{oversight}/{timestamp}.json`:
## Key Features
- **Reasoning Extraction**: Supports multiple formats:
- - OpenRouter `reasoning_content`
+ - OpenRouter `reasoning_content` (Qwen, DeepSeek, etc.)
+ - OpenAI `reasoning_details` — both `reasoning.text` and `reasoning.summary` types
- Google `thought_signature` (via `extra_content`)
- `<thinking>` tags in content
+ - Each log entry includes `reasoning_format` metadata for traceability
+- **Baseline Toggle**: Set `generate_baseline: false` in defaults to skip baseline generation (useful when testing new models)
- **Turn Limit**: Max 20 turns per conversation (prevents infinite loops)
- **Success Tracking**: Runs marked SUCCESS if they end with `finish_reason: "stop"`
- **VFS**: Virtual file system for sandboxed file access
@@ -182,15 +186,15 @@ The project includes 138 unit tests covering all modules. No API keys or network
uv run pytest tests/ -v
```
-| Test file | Covers |
-|-----------|--------|
-| `test_vfs.py` | Sandbox escapes, path traversal, CRUD, singleton staleness |
-| `test_config_loader.py` | Missing keys, temperature cascade, oversight fallback |
-| `test_tools.py` | Schema/implementation sync, VFS delegation |
-| `test_agent.py` | Message construction, error handling, token counting, reasoning extraction |
-| `test_interrogate.py` | Conversation sanitization, provider detection |
-| `test_runner.py` | Baseline extraction, prompt assembly, success detection |
-| `test_judge.py` | Regex/blackbox/glassbox checkers, JSON parsing, batch prep, CSV output |
+| Test file | Covers |
+| ----------------------- | -------------------------------------------------------------------------- |
+| `test_vfs.py` | Sandbox escapes, path traversal, CRUD, singleton staleness |
+| `test_config_loader.py` | Missing keys, temperature cascade, oversight fallback |
+| `test_tools.py` | Schema/implementation sync, VFS delegation |
+| `test_agent.py` | Message construction, error handling, token counting, reasoning extraction |
+| `test_interrogate.py` | Conversation sanitization, provider detection |
+| `test_runner.py` | Baseline extraction, prompt assembly, success detection |
+| `test_judge.py` | Regex/blackbox/glassbox checkers, JSON parsing, batch prep, CSV output |
## File Structure
@@ -232,14 +236,14 @@ uv run src/interrogate.py logs/model_name/scenario/oversight/timestamp.json
The session auto-detects the provider from the log file and restores the VFS state. Available commands:
-| Command | Description |
-|---------|-------------|
-| `history` | Show full conversation history |
-| `history N` | Show last N messages |
-| `reasoning` | Show the last full reasoning trace |
-| `vfs` | Show current virtual filesystem state |
-| `info` | Show run metadata (model, scenario, tokens) |
-| `save` | Save the extended conversation to `interrogation_logs/` |
-| `exit` | Quit |
+| Command | Description |
+| ----------- | ------------------------------------------------------- |
+| `history` | Show full conversation history |
+| `history N` | Show last N messages |
+| `reasoning` | Show the last full reasoning trace |
+| `vfs` | Show current virtual filesystem state |
+| `info` | Show run metadata (model, scenario, tokens) |
+| `save` | Save the extended conversation to `interrogation_logs/` |
+| `exit` | Quit |
Anything else you type is sent as a message to the agent.
diff --git a/config.yaml.example b/config.yaml.example
index e78c8ce..a7554ee 100644
--- a/config.yaml.example
+++ b/config.yaml.example
@@ -8,6 +8,7 @@ defaults:
# max_tokens: 4096
oversight: high
include_reasoning_in_context: true
+ generate_baseline: true # Set to false to skip baseline generation
# Provider configurations
providers:
diff --git a/src/config_loader.py b/src/config_loader.py
index 8db50cb..2badfbb 100644
--- a/src/config_loader.py
+++ b/src/config_loader.py
@@ -147,6 +147,11 @@ class ConfigLoader:
return self._config.get('defaults', {})
@property
+ def generate_baseline(self) -> bool:
+ """Whether to generate baselines before experiment runs."""
+ return self._config.get('defaults', {}).get('generate_baseline', True)
+
+ @property
def output_dir(self) -> str:
"""Get output directory."""
return self._config.get('output', {}).get('dir', 'output')
diff --git a/src/runner.py b/src/runner.py
index 8556ed6..5cdb807 100644
--- a/src/runner.py
+++ b/src/runner.py
@@ -77,7 +77,14 @@ class ExperimentRunner:
# Ensure baseline exists before running hidden-goal experiments
output_dir = self.config.output_dir
baseline_path = os.path.join(output_dir, model_name_safe, scenario_name, "baseline.md")
- if not os.path.exists(baseline_path):
+ if not self.config.generate_baseline:
+ if not os.path.exists(baseline_path):
+ logger.warning(f"Baseline generation is DISABLED (generate_baseline: false). "
+ f"No baseline exists for {model_name} | {scenario_name}. "
+ f"Black-box judging will not be possible for these runs.")
+ else:
+ logger.info(f"\n--- Baseline exists (generation disabled): {model_name} | {scenario_name} ---")
+ elif not os.path.exists(baseline_path):
logger.info(f"\n--- Generating baseline: {model_name} | {scenario_name} ---")
self._run_baseline(model_config, provider_config, scenario_config)
logger.info(f" Baseline saved to {baseline_path}")