summaryrefslogtreecommitdiff
path: root/notes/v2_redesign/05_dissertation_pivot.md
diff options
context:
space:
mode:
Diffstat (limited to 'notes/v2_redesign/05_dissertation_pivot.md')
-rw-r--r--notes/v2_redesign/05_dissertation_pivot.md54
1 files changed, 54 insertions, 0 deletions
diff --git a/notes/v2_redesign/05_dissertation_pivot.md b/notes/v2_redesign/05_dissertation_pivot.md
new file mode 100644
index 0000000..d6a3729
--- /dev/null
+++ b/notes/v2_redesign/05_dissertation_pivot.md
@@ -0,0 +1,54 @@
+# Dissertation Pivot — April 2026
+
+**Date:** 2026-04-13
+**Deadline:** End of April 2026 (dissertation submission)
+
+---
+
+## Revised Plan
+
+**Dissertation:** Submit with v1 as primary empirical contribution. v1 is complete, analysed, and ready to write up.
+
+**Publication (separate):** Complete v2 experiments (Study 1 + Study 2) and publish findings independently after dissertation submission.
+
+---
+
+## Dissertation Structure (v1-based)
+
+### Empirical Content
+- v1 experiment: 420 runs, 68.3% overall deception rate
+- Key findings:
+ - Oversight reduces deception frequency (74.3% → 57.9%) but not sophistication (69.3% STRATEGIC across all levels)
+ - Scenario effect: child_protection (87%) vs corporate_sabotage (50%) — noted as confounded, motivates v2
+ - Model variation: Gemini 3.1 Pro (100%) to GLM 4.7 Flash (40%)
+
+### Methodology
+- Experiment design: 7 models × 2 scenarios × 3 oversight levels × 10 runs
+- Judge system: glassbox + blackbox + regex validation
+- Three-layer analysis strategy (pooled, per-model, sensitivity)
+
+### v2 as Future Work
+- Acknowledge v1 limitations (confounded scenario comparison, insufficient runs per cell)
+- Propose two-phase design (Study 1: oversight effect, Study 2: rationalization mechanism)
+- Pre-registered analysis: Cochran-Armitage trend test for oversight as ordinal predictor
+
+---
+
+## Timeline
+
+| Task | Status | Deadline |
+|------|--------|----------|
+| Dissertation writing | Not started | ~2026-04-30 |
+| Dissertation submission | Not submitted | 2026-04-30 |
+| Judge validation pilot | Pending | Post-dissertation |
+| Study 1 (270 runs) | Pending | Post-dissertation |
+| Study 2 (270 runs) | Pending | Post-dissertation |
+| Publication submission | Pending | Post-dissertation |
+
+---
+
+## Rationale
+
+Starting v2 experiments without validated judges and insufficient time to complete analysis and writing is high risk. v1 alone is sufficient for a complete dissertation — the research question is well-motivated, the methodology is sound, and the findings are substantive (oversight × sophistication independence is a non-obvious result worth publishing on its own).
+
+v2 experiments and publication are deferred to after dissertation submission, when adequate time is available for proper judge validation, execution, and analysis.