diff options
| author | CaptainJack2491 <jayrupnakawala@gmail.com> | 2026-04-15 21:18:57 +0100 |
|---|---|---|
| committer | CaptainJack2491 <jayrupnakawala@gmail.com> | 2026-04-15 21:18:57 +0100 |
| commit | 6cc92304379996f68bb2df96a5e4b637b4c4804b (patch) | |
| tree | 6cdbc8b5484934927cd6ac989e91fef44a60bb4c /notes | |
| parent | d25f8fc2e2b02c04489ca1960b83ff7a70e7317a (diff) | |
Study 1: 270 runs complete + judge validation pipeline
- config_study1.yaml: 3 models × 3 oversight × bare framing × n=30
- Judge validation: 54-run subset, gold=Sonnet 4.6, proxy=Grok 4.1 Fast (BB κ=0.702) + GPT-4.1 (GB κ=0.878)
- Fix OpenAI batch provider: BytesIO, method/url fields, response.body parsing
- Scripts: extract_subset.py, judge_validation.py
- Dissertation chapters updated (intro, methodology, results, conclusion)
Diffstat (limited to 'notes')
| -rw-r--r-- | notes/v2_redesign/06_2026-04-13_cleanup_and_reorganisation.md | 163 | ||||
| -rw-r--r-- | notes/v2_redesign/07_judge_validation_results.md | 52 |
2 files changed, 215 insertions, 0 deletions
diff --git a/notes/v2_redesign/06_2026-04-13_cleanup_and_reorganisation.md b/notes/v2_redesign/06_2026-04-13_cleanup_and_reorganisation.md new file mode 100644 index 0000000..6c1b95c --- /dev/null +++ b/notes/v2_redesign/06_2026-04-13_cleanup_and_reorganisation.md @@ -0,0 +1,163 @@ +# Cleanup and Repository Reorganisation + +**Date:** 2026-04-13 +**Purpose:** Document repository cleanup actions taken before dissertation submission + +--- + +## Deleted Files (Garbage Removed) + +| File | Reason | +|------|--------| +| AGENT.md | Stale agent configuration | +| PLAN_judge_dashboard.md | Ad-hoc planning doc | +| PLAN_judge_multi_model.md | Ad-hoc planning doc | +| analyze.py | Debug script | +| analyze_gpt.py | Debug script | +| bruh.csv | Debug artifact | +| config_test_eval.yaml | Redundant config | +| grok_batchapi.md | Stray note | +| oversight_competence_filter.md | Stray note | +| pharma_trial_scenario_spec.md | Stray spec | +| pilot_judge.py | Debug script | +| results_plan.md | Stray planning doc | +| src/quick_judge.py | Debug script | +| task.md | Stray task doc | +| test.config.yaml | Debug config | +| test.csv | Debug artifact | + +## Updated Files + +### .gitignore +- Added `videos/` to gitignore (656MB recordings, already gitignored locally) + +## Commits Made (11 Total) + +### Commit 1: ignore +``` +ignore: add videos/ to gitignore +``` +- Updated `.gitignore` to exclude `videos/` + +### Commit 2: docs +``` +docs: update dissertation intro, lit review, and methodology chapters +``` +- `docs/final/intro.typ` +- `docs/final/lit.typ` +- `docs/final/methodology.typ` +- `docs/final/methodology_1.typ` + +### Commit 3: viz +``` +viz: add architecture diagrams for thesis +``` +- `viz/ai_evaluation_dashboard.png` +- `viz/ai_evaluation_heatmaps.png` +- `viz/ai_evaluation_oversight_analysis.png` +- `viz/ai_evaluation_radar.png` +- `viz/architecture.mmd` +- `viz/architecture.pdf` +- `viz/architecture.svg` +- `viz/gem.png` +- `viz/model_evaluation_summary.csv` + +### Commit 4: papers +``` +papers: add 2512.16041v1 deception alignment paper +``` +- `papers/2512.16041v1.pdf` + +### Commit 5: docs/paper +``` +docs: add paper draft with alternative structure for publication +``` +- `docs/paper/TODO.md` +- `docs/paper/abstract.typ` +- `docs/paper/conclusion.typ` +- `docs/paper/introduction.typ` +- `docs/paper/main.pdf` +- `docs/paper/main.typ` +- `docs/paper/methodology.typ` +- `docs/paper/plan.md` +- `docs/paper/questions.md` +- `docs/paper/references.bib` +- `docs/paper/results.md` +- `docs/paper/results.typ` +- `docs/paper/uel.svg` + +### Commit 6: judge_validation +``` +judge_validation: add v2 judge comparison results +``` +- `judge_validation/comparison_results.csv` +- `judge_validation/results_claude_haiku_4_5_20251001.csv` +- `judge_validation/results_claude_sonnet_4_6.csv` +- `judge_validation/results_gpt_4_1_mini.csv` +- `judge_validation/results_gpt_5_mini_2025_08_07.csv` +- `judge_validation/results_grok_4_1_fast_non_reasoning.csv` +- `judge_validation/results_grok_4_1_fast_reasoning.csv` + +### Commit 7: docs +``` +docs: add judging specification and config template +``` +- `judging_spec.md` (kept for documentation value) +- `example.config.yaml` (template for other users, main config.yaml is gitignored) + +### Commit 8: notes +``` +notes: add dissertation pivot plan and additional research notes +``` +- `notes/concerns.md` +- `notes/safe_goal_experiments.md` +- `notes/v2_redesign/05_dissertation_pivot.md` + +### Commit 9: docs +``` +docs: add progress presentation from start of Term 2 (Jan 2026) +``` +- `docs/04_ppt/Presentation Template.pdf` +- `docs/04_ppt/harward.csl` +- `docs/04_ppt/main.html` +- `docs/04_ppt/main.qmd` +- `docs/04_ppt/plan.md` +- `docs/04_ppt/progress_presentation.html` + +### Commit 10: api +``` +api: add chat interface styles to results dashboard +``` +- `api/static/css/results.css` (152 insertions) +- `api/static/js/ui.js` (1 insertion, 1 deletion) + +### Commit 11: deps +``` +deps: add scikit-learn for statistical analysis +``` +- `pyproject.toml` (added scikit-learn>=1.8.0) +- `uv.lock` + +## Retained (Not Committed) + +### Untracked - Keep as-is +- `save.csv`, `save_if.csv` — debug artifacts, keep untracked +- `docs/final/claude_mathodology.md` — older draft methodology +- `docs/final/main1.pdf/typ` — older draft version +- `docs/final/methodology_1.pdf` — alternate methodology draft +- `docs/final/template.typ`, `harward.csl`, `uel.svg` — reference files + +### Modified - Active Work (Left Uncommitted) +- `src/judge/batch_providers.py` — v2 judge implementation +- `src/judge/judge.py` — v2 judge system +- `src/judge/prompts.py` — v2 judge prompts +- `scripts/judge_comparison.py` — judge utility +- `scripts/test_dashboard_ui.py` — dashboard utility +- `config_dryrun.yaml` — dry run config (intentionally uncommitted) +- `logs/` — separate worktree on different branch + +## Final Status + +- Branch: `main`, 24 commits ahead of `origin/main` +- Clean state for dissertation work +- v2 judge system and experiment code left in working state for post-dissertation completion diff --git a/notes/v2_redesign/07_judge_validation_results.md b/notes/v2_redesign/07_judge_validation_results.md new file mode 100644 index 0000000..faf32de --- /dev/null +++ b/notes/v2_redesign/07_judge_validation_results.md @@ -0,0 +1,52 @@ +# Judge Validation Results — 2026-04-15 + +## Study 1 Subset + +- 54 runs extracted from Study 1 (20% stratified subset out of 270, seed=42) +- Logs: `logs/v2_study1_judge_subset/` + +## Gold Standard + +- Blackbox: `claude-sonnet-4-6` (Anthropic) +- Glassbox: `claude-sonnet-4-6` (Anthropic) + +## Results + +| Model | BB κ | GB κ | Sph κ | Pass | +| --------------------------- | --------- | --------- | --------- | ------- | +| gpt-4.1 | 0.351 | 0.878 | 0.538 | no | +| gpt-4.1-mini | 0.202 | 0.834 | 0.345 | no | +| **grok-4-1-fast-reasoning** | **0.702** | **0.812** | **0.108** | **YES** | +| grok-4.20-reasoning | 0.549 | 0.919 | 0.198 | no | +| grok-4.20-non-reasoning | 0.390 | 0.878 | 0.117 | no | + +## Selected Proxy Judges + +Split configuration — best model per prong: + +| Prong | Proxy Judge | κ vs Gold | Rationale | +| ------------ | ----------------------------- | ---------------------------- | ------------------------------- | +| **Blackbox** | grok-4-1-fast-reasoning (xAI) | 0.702 | Only model to pass BB threshold | +| **Glassbox** | gpt-4.1 (OpenAI) | 0.878 (intent), 0.538 (soph) | Best sophistication agreement | + +### Config + +```yaml +judge: + blackbox: + model: grok-4-1-fast-reasoning + provider: xai + temperature: 0 + glassbox: + model: gpt-4.1 + provider: openai + temperature: 0 +``` + +## Notes + +- All models show high glassbox agreement (κ > 0.8) — intent classification is consistent across judges +- Blackbox agreement is more variable — output framing/categorization is harder to align on +- Sophistication agreement is low across the board — tier classification is noisy (flagged as limitation) +- Three different provider families used: Anthropic (gold), xAI (BB proxy), OpenAI (GB proxy) — no same-family bias +- Gold standard used a single model (Claude Sonnet 4.6) for both prongs for consistency |
