1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
|
# Web GUI Implementation Plan
## Overview
A lightweight, decoupled web interface for the AI Agent Reasoning Experiment Framework. The goal is to allow non-technical users (e.g., supervisors, dissertation examiners) to understand and interact with the project without using the CLI.
**Core Principle**: Keep it fast, reactive, and modular. No bloat. The GUI is a separate layer that wraps the existing core project without modifying it.
---
## Architecture
### Decoupled Design
```
┌─────────────────────────────────────────────────────────────┐
│ Web GUI (Frontend) │
│ Vanilla JS + HTML + CSS │
└────────────────────────────┬────────────────────────────────┘
│ HTTP / SSE
┌────────────────────────────▼────────────────────────────────┐
│ FastAPI Backend (API) │
│ api/server.py (new, separate) │
│ - Wraps core functions (ConfigLoader, run_from_config) │
│ - Exposes endpoints for scenarios, models, runs, logs │
└────────────────────────────┬────────────────────────────────┘
│
┌────────────────────────────▼────────────────────────────────┐
│ Core Project (Unchanged) │
│ - src/main.py, src/judge_runner.py │
│ - config.yaml, scenarios/, logs/ │
└─────────────────────────────────────────────────────────────┘
```
### Why This Approach?
1. **Zero Changes to Core**: The dissertation code remains untouched and pure.
2. **Modular**: Adding new scenarios automatically updates the GUI.
3. **No Bloat**: FastAPI + Vanilla JS. Fast, reactive, minimal dependencies.
4. **Showcase-Ready**: Clean dashboard for non-technical users.
---
## File Structure
```
dissertation/
├── api/ # New API layer
│ ├── __init__.py
│ ├── server.py # FastAPI application
│ ├── endpoints/
│ │ ├── __init__.py
│ │ ├── config.py # Config read/write endpoints
│ │ ├── scenarios.py # Scenario discovery endpoints
│ │ ├── runs.py # Experiment run triggers
│ │ └── results.py # Results/logs fetching
│ └── static/ # Frontend (served by FastAPI)
│ ├── index.html
│ ├── app.js
│ ├── style.css
│ └── assets/
│
├── src/ # Core project (UNCHANGED)
│ ├── main.py
│ ├── runner.py
│ ├── config_loader.py
│ └── ...
│
├── docs/
│ └── web_gui_plan.md # This file
│
└── config.yaml
```
---
## API Endpoints
### 1. Configuration
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/api/config` | Read current `config.yaml` |
| PUT | `/api/config` | Update/save `config.yaml` |
| GET | `/api/logging` | Get current logging configuration (file path, level) |
### 2. Discovery
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/api/scenarios` | List all available scenarios from `scenarios/` |
| GET | `/api/scenarios/{name}` | Get details (oversight levels, files) for a scenario |
| GET | `/api/models` | List models from config |
| GET | `/api/providers` | List providers from config |
### 3. Execution
| Method | Endpoint | Description |
|--------|----------|-------------|
| POST | `/api/run` | Start an experiment run (non-blocking) |
| GET | `/api/run/status` | Get current run status (idle/running/complete) |
| DELETE | `/api/run` | Cancel current run |
| GET | `/api/logs/stream` | Server-Sent Events (SSE) for live log streaming |
### 4. Results
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/api/results` | Get experiment results (CSV data as JSON) |
| GET | `/api/results/images` | List generated visualization images |
| GET | `/api/results/images/{name}` | Serve a specific image |
| GET | `/api/judge/results` | Get judge results if available |
---
## Implementation Steps
### Phase 1: Backend (FastAPI)
1. **Setup**: Add `fastapi`, `uvicorn` to dependencies (via `uv add`)
2. **Server**: Create `api/server.py` with basic FastAPI app
3. **Config Endpoint**: Wrap `ConfigLoader` to expose config read/write
4. **Scenario Discovery**: Auto-scan `scenarios/` directory
5. **Run Trigger**: Execute `uv run src/main.py` as subprocess
6. **Log Streaming**: Implement SSE endpoint to tail log file
7. **Results Endpoint**: Read CSV files and serve as JSON
### Phase 2: Frontend (Vanilla JS)
1. **Static Files**: Create `api/static/` directory
2. **index.html**: Main dashboard layout
3. **style.css**: Clean, minimal styling
4. **app.js**:
- Fetch scenarios/models and populate dropdowns
- Handle "Run" button click → POST to `/api/run`
- Connect to `/api/logs/stream` for live output
- Display results in tables
- Render images from `/api/results/images`
### Phase 3: Integration & Polish
1. **Auto-discovery**: Ensure new scenarios appear automatically
2. **Error Handling**: Graceful errors for missing API keys, etc.
3. **Run Status**: Visual indicator (spinner/green check) during runs
---
## Key Implementation Details
### Log File Location
The log file path is **not hardcoded**. It is read dynamically from `config.yaml` at runtime:
```python
# From config_loader.py
logging_config = config.logging_config
log_file = logging_config.get('file') # e.g., 'logs/gpt/experiment.log'
```
The API uses this to tail the correct file during live streaming.
### Non-Blocking Runs
Experiment runs can take a long time. The API uses `asyncio` or `subprocess.Popen` to start the run in the background, returning a `run_id` or status immediately. The frontend can poll for status or stream logs in real-time.
### Server-Sent Events (SSE) for Log Streaming
Instead of WebSockets (which add bloat), we use SSE:
```python
@app.get("/api/logs/stream")
async def log_stream():
async def event_generator():
# Tail log file line by line
yield {"event": "log", "data": line}
return EventSourceResponse(event_generator())
```
---
## Dependencies
New dependencies to add via `uv add`:
```toml
[dependencies]
fastapi = ">=0.115.0"
uvicorn = {extras = ["standard"], version = ">=0.32.0"}
sse-starlette = ">=2.0.0"
```
Frontend uses **zero** external dependencies (pure Vanilla JS).
---
## Running the GUI
```bash
# Start the API server
uv run api/server.py
# Open in browser
# http://localhost:8000
```
---
## Future Considerations (Out of Scope for Initial Version)
- [ ] Authentication (not needed for local dissertation showcase)
- [ ] Multi-user support
- [ ] Judge runner integration via GUI
- [ ] Dark mode
- [ ] Mobile responsiveness
---
## Author
Created: March 2026
Purpose: Dissertation showcase for AI Agent Reasoning Experiment Framework
|