summaryrefslogtreecommitdiff
path: root/README.md
blob: 92ca61e580c30a7c7101e180911d511d947b9095 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
# AI Agent Reasoning Experiment Framework

A framework for running AI agent experiments with structured reasoning extraction, logging, and analysis.

## What It Does

- Runs AI agents with tool-calling capabilities (file system access)
- Extracts and displays reasoning/thought processes from models
- Logs conversations with metadata (token counts, finish reasons, turn counts)
- Supports multiple model providers via OpenAI-compatible APIs
- Configurable via YAML

## Quick Start

```bash
# Install dependencies
uv sync

# Run all experiments from config.yaml
uv run src/main.py

# Run a specific scenario
uv run src/main.py --scenario scenarios/corporate_sabotage_v2 --oversight high
```

## Configuration (`config.yaml`)

```yaml
providers:
  openai:
    api_key_env: OPENAI_API_KEY
    base_url: https://api.openai.com/v1
  google:
    api_key_env: GOOGLE_API_KEY
    base_url: https://generativelanguage.googleapis.com/v1beta/openai/
  openrouter:
    api_key_env: OPENROUTER_API_KEY
    base_url: https://openrouter.ai/api/v1

models:
  - id: gemini-3-flash-preview
    provider: google
    temperature: 1
    extra_body:
      google:
        thinking_config:
          include_thoughts: True

scenarios:
  - path: scenarios/corporate_sabotage_v2
    runs: 1

oversight_levels:
  - high

output:
  dir: logs
  save_vfs_state: false  # Set true to include final VFS state in logs
```

## Scenario Structure

```
scenarios/
  scenario_name/
    user.md           # User prompt
    benign_goal.md    # Visible system prompt
    hidden_goal.md    # Hidden system prompt (for experiments)
    data/             # Files accessible to agent
      file1.txt
      file2.txt
```

## Output Format

Logs are saved to `logs/{model}/{scenario}/{oversight}/{timestamp}.json`:

```json
{
  "run_id": "...",
  "model": "gemini-3-flash-preview",
  "scenario": "corporate_sabotage_v2",
  "oversight_level": "high",
  "total_tokens": 1500,
  "prompt_tokens": 1000,
  "completion_tokens": 500,
  "conversation": [
    {"role": "system", "content": "..."},
    {"role": "user", "content": "..."},
    {
      "role": "assistant",
      "reasoning": "...",
      "content": "...",
      "tool_calls": [...],
      "finish_reason": "stop",
      "turn_count": 1
    }
  ]
}
```

## Key Features

- **Reasoning Extraction**: Supports multiple formats:
  - OpenRouter `reasoning_content`
  - Google `thought_signature` (via `extra_content`)
  - `<thinking>` tags in content
- **Turn Limit**: Max 20 turns per conversation (prevents infinite loops)
- **Success Tracking**: Runs marked SUCCESS if they end with `finish_reason: "stop"`
- **VFS**: Virtual file system for sandboxed file access

## File Structure

```
src/
  agent.py       # Main agent logic, OpenAI SDK integration
  config_loader.py  # YAML config parsing
  main.py        # Entry point
  runner.py      # Experiment orchestration
  tools.py       # Available tools (list_files, read_file, etc.)
  vfs.py         # Virtual file system
scenarios/       # Scenario definitions
logs/            # Output logs
```

## API Keys

Set API keys via environment variables (or `.env` file):

```bash
export OPENAI_API_KEY="..."
export GOOGLE_API_KEY="..."
export OPENROUTER_API_KEY="..."
```

## Interrogation [Not implimented correctly yet]

Replay and continue conversations from logs:

```bash
uv run src/interrogate.py logs/gemini-3-flash-preview/...
```