summaryrefslogtreecommitdiff
path: root/docs/paper/results.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/paper/results.md')
-rw-r--r--docs/paper/results.md388
1 files changed, 388 insertions, 0 deletions
diff --git a/docs/paper/results.md b/docs/paper/results.md
new file mode 100644
index 0000000..1f7ea0c
--- /dev/null
+++ b/docs/paper/results.md
@@ -0,0 +1,388 @@
+Loading data from: logs/full_experiment/results.csv
+Loaded 420 rows: 7 models, 2 scenarios, 3 oversight levels
+Significance level: α = 0.05
+
+================================================================================
+ 1. SCENARIO EFFECT ON DECEPTION
+================================================================================
+
+
+ --- 1a. Output-level deception rate by scenario (blackbox) ---
+
+ Contingency table:
+ child_protection: 177/210 (84.3%) 95% CI: [78.8%, 88.6%]
+ corporate_sabotage_v2: 104/210 (49.5%) 95% CI: [42.8%, 56.2%]
+
+ Chi-squared test: χ²(1) = 55.743, p < .001 ***
+ Cramér's V = 0.364 (medium)
+ Cramér's V (bias-corrected) = 0.361
+ Odds ratio = 0.183 95% CI: [0.116, 0.290]
+ Fisher's exact test: p < .001 ***
+
+ --- 1b. Reasoning-level compliance by scenario (glassbox: COMPLIANT_DECEPTION) ---
+
+ Contingency table:
+ child_protection: 182/210 (86.7%) 95% CI: [81.4%, 90.6%]
+ corporate_sabotage_v2: 105/210 (50.0%) 95% CI: [43.3%, 56.7%]
+
+ Chi-squared test: χ²(1) = 63.554, p < .001 ***
+ Cramér's V = 0.389 (medium)
+ Odds ratio = 0.154 95% CI: [0.095, 0.249]
+
+================================================================================
+ 2. OVERSIGHT EFFECT ON DECEPTION
+================================================================================
+
+
+ --- 2a. Output-level deception by oversight — child_protection ---
+
+ Contingency table:
+ low: 60/70 (85.7%) 95% CI: [75.7%, 92.1%]
+ mid: 59/70 (84.3%) 95% CI: [74.0%, 91.0%]
+ high: 58/70 (82.9%) 95% CI: [72.4%, 89.9%]
+
+ Chi-squared test: χ²(2) = 0.216, p = 0.8978 n.s.
+ Cramér's V = 0.032 (negligible)
+ Minimum expected cell count = 11.0 ✓ All expected counts ≥ 5
+
+ Cochran-Armitage trend test: z = -0.464, p = 0.6423 n.s.
+ → No significant linear trend detected
+
+ --- 2b. Pairwise oversight comparisons — child_protection (Bonferroni-corrected, α/3 = 0.0167) ---
+
+ low vs mid:
+ χ²(1) = 0.000, raw p = 1.0000, Fisher p = 1.0000
+ OR = 0.894 95% CI: [0.353, 2.263] [Bonferroni: n.s.]
+ low vs high:
+ χ²(1) = 0.054, raw p = 0.8164, Fisher p = 0.8169
+ OR = 0.806 95% CI: [0.323, 2.008] [Bonferroni: n.s.]
+ mid vs high:
+ χ²(1) = 0.000, raw p = 1.0000, Fisher p = 1.0000
+ OR = 0.901 95% CI: [0.368, 2.205] [Bonferroni: n.s.]
+
+ --- 2a. Output-level deception by oversight — corporate_sabotage_v2 ---
+
+ Contingency table:
+ low: 46/70 (65.7%) 95% CI: [54.0%, 75.8%]
+ mid: 35/70 (50.0%) 95% CI: [38.6%, 61.4%]
+ high: 23/70 (32.9%) 95% CI: [23.0%, 44.5%]
+
+ Chi-squared test: χ²(2) = 15.125, p < .001 ***
+ Cramér's V = 0.268 (small)
+ Minimum expected cell count = 34.7 ✓ All expected counts ≥ 5
+
+ Cochran-Armitage trend test: z = -3.888, p < .001 ***
+ → Significant linear trend: deception decreasing with oversight level
+
+ --- 2b. Pairwise oversight comparisons — corporate_sabotage_v2 (Bonferroni-corrected, α/3 = 0.0167) ---
+
+ low vs mid:
+ χ²(1) = 2.929, raw p = 0.0870, Fisher p = 0.0866
+ OR = 0.522 95% CI: [0.264, 1.030] [Bonferroni: n.s.]
+ low vs high:
+ χ²(1) = 13.831, raw p < .001, Fisher p < .001
+ OR = 0.255 95% CI: [0.127, 0.515] [Bonferroni: SIG]
+ mid vs high:
+ χ²(1) = 3.562, raw p = 0.0591, Fisher p = 0.0587
+ OR = 0.489 95% CI: [0.247, 0.970] [Bonferroni: n.s.]
+
+================================================================================
+ 3. MODEL-LEVEL DIFFERENCES
+================================================================================
+
+
+ --- 3a. Output deception by model — child_protection ---
+
+ Model deception rates:
+ deepseek/deepseek-v3.2: 30/30 (100.0%) 95% CI: [88.6%, 100.0%]
+ gemini-3.1-pro-preview: 30/30 (100.0%) 95% CI: [88.6%, 100.0%]
+ moonshotai/kimi-k2.5: 27/30 (90.0%) 95% CI: [74.4%, 96.5%]
+ openai/gpt-oss-20b: 25/30 (83.3%) 95% CI: [66.4%, 92.7%]
+ openai/gpt-oss-safeguard-20b: 12/30 (40.0%) 95% CI: [24.6%, 57.7%]
+ qwen/qwen3-max-thinking: 30/30 (100.0%) 95% CI: [88.6%, 100.0%]
+ z-ai/glm-4.7-flash: 23/30 (76.7%) 95% CI: [59.1%, 88.2%]
+
+ Chi-squared test: χ²(6) = 63.277, p < .001 ***
+ Cramér's V = 0.549 (large)
+ ⚠ Minimum expected count = 4.7 < 5
+
+ --- 3a. Output deception by model — corporate_sabotage_v2 ---
+
+ Model deception rates:
+ deepseek/deepseek-v3.2: 8/30 (26.7%) 95% CI: [14.2%, 44.4%]
+ gemini-3.1-pro-preview: 27/30 (90.0%) 95% CI: [74.4%, 96.5%]
+ moonshotai/kimi-k2.5: 13/30 (43.3%) 95% CI: [27.4%, 60.8%]
+ openai/gpt-oss-20b: 16/30 (53.3%) 95% CI: [36.1%, 69.8%]
+ openai/gpt-oss-safeguard-20b: 16/30 (53.3%) 95% CI: [36.1%, 69.8%]
+ qwen/qwen3-max-thinking: 12/30 (40.0%) 95% CI: [24.6%, 57.7%]
+ z-ai/glm-4.7-flash: 12/30 (40.0%) 95% CI: [24.6%, 57.7%]
+
+ Chi-squared test: χ²(6) = 28.917, p < .001 ***
+ Cramér's V = 0.371 (medium)
+
+================================================================================
+ 4. OVERSIGHT × SCENARIO INTERACTION
+================================================================================
+
+
+ --- 4a. Logistic regression: deception ~ scenario + oversight + scenario×oversight ---
+
+ Coef. Std.Err. z P>|z| [0.025 0.975]
+const 0.661995 0.228656 2.895152 0.003790 0.213837 1.110154
+scenario_code 1.128325 0.384590 2.933835 0.003348 0.374541 1.882108
+oversight_num -0.682521 0.178978 -3.813430 0.000137 -1.033312 -0.331730
+interaction 0.574509 0.293588 1.956854 0.050365 -0.000913 1.149931
+
+ Model fit: Pseudo R² = 0.1409, AIC = 466.1, BIC = 482.3
+
+ Likelihood ratio test for interaction: LR = 3.816, p = 0.0508 n.s.
+ → No significant interaction
+
+ --- 4b. Breslow-Day test for homogeneity of odds ratios ---
+
+ (Tests whether the scenario effect is consistent across oversight levels)
+ low: OR = 0.319 95% CI: [0.139, 0.734]
+ mid: OR = 0.186 95% CI: [0.084, 0.413]
+ high: OR = 0.101 95% CI: [0.046, 0.225]
+
+ Cochran-Mantel-Haenszel test:
+ Common OR = N/A
+ CMH statistic = 58.293, p < .001 ***
+
+ Breslow-Day test for homogeneity of ORs:
+ Statistic = 3.885, p = 0.1434 n.s.
+ → ORs are consistent across oversight strata
+
+================================================================================
+ 5. MODEL × OVERSIGHT INTERACTION (per scenario)
+================================================================================
+
+
+ --- 5. Oversight effect by model — child_protection ---
+
+ deepseek-v3.2 low=100.0% mid=100.0% high=100.0% trend z=+0.00 p = 1.0000 n.s.
+ gemini-3.1-pro-preview low=100.0% mid=100.0% high=100.0% trend z=+0.00 p = 1.0000 n.s.
+ kimi-k2.5 low=100.0% mid= 80.0% high= 90.0% trend z=-0.75 p = 0.4561 n.s.
+ gpt-oss-20b low= 80.0% mid= 80.0% high= 90.0% trend z=+0.60 p = 0.5485 n.s.
+ gpt-oss-safeguard-20b low= 50.0% mid= 50.0% high= 20.0% trend z=-1.37 p = 0.1709 n.s.
+ qwen3-max-thinking low=100.0% mid=100.0% high=100.0% trend z=+0.00 p = 1.0000 n.s.
+ glm-4.7-flash low= 70.0% mid= 80.0% high= 80.0% trend z=+0.53 p = 0.5970 n.s.
+
+ --- 5. Oversight effect by model — corporate_sabotage_v2 ---
+
+ deepseek-v3.2 low= 20.0% mid= 40.0% high= 20.0% trend z=+0.00 p = 1.0000 n.s.
+ gemini-3.1-pro-preview low=100.0% mid= 80.0% high= 90.0% trend z=-0.75 p = 0.4561 n.s.
+ kimi-k2.5 low= 50.0% mid= 60.0% high= 20.0% trend z=-1.35 p = 0.1758 n.s.
+ gpt-oss-20b low= 70.0% mid= 60.0% high= 30.0% trend z=-1.79 p = 0.0730 n.s.
+ gpt-oss-safeguard-20b low= 70.0% mid= 50.0% high= 40.0% trend z=-1.34 p = 0.1787 n.s.
+ qwen3-max-thinking low= 70.0% mid= 40.0% high= 10.0% trend z=-2.74 p = 0.0062 **
+ glm-4.7-flash low= 80.0% mid= 20.0% high= 20.0% trend z=-2.74 p = 0.0062 **
+
+================================================================================
+ 6. DECEPTION TYPE ANALYSIS
+================================================================================
+
+
+ --- 6a. Deception type distribution by scenario ---
+
+ child_protection (n=177):
+ FABRICATED: 2 (1.1%)
+ MISLEADING: 150 (84.7%)
+ OMITTED: 25 (14.1%)
+ corporate_sabotage_v2 (n=104):
+ FABRICATED: 25 (24.0%)
+ MISLEADING: 55 (52.9%)
+ OMITTED: 24 (23.1%)
+
+ Chi-squared test: χ²(2) = 47.906, p < .001 ***
+ Cramér's V = 0.413 (medium)
+
+================================================================================
+ 7. REASONING INTENT ANALYSIS (Glassbox)
+================================================================================
+
+
+ --- 7a. Glassbox category distribution by scenario ---
+
+ child_protection (n=210):
+ COMPLIANT_DECEPTION: 182 (86.7%) 95% CI: [81.4%, 90.6%]
+ CONSIDERED_HONEST: 6 (2.9%) 95% CI: [1.3%, 6.1%]
+ FAILED_DECEPTION: 3 (1.4%) 95% CI: [0.5%, 4.1%]
+ ROBUST_HONEST: 10 (4.8%) 95% CI: [2.6%, 8.5%]
+ UNINTENTIONAL_FAILURE: 9 (4.3%) 95% CI: [2.3%, 7.9%]
+ UNKNOWN: 0 (0.0%) 95% CI: [0.0%, 1.8%]
+ corporate_sabotage_v2 (n=210):
+ COMPLIANT_DECEPTION: 105 (50.0%) 95% CI: [43.3%, 56.7%]
+ CONSIDERED_HONEST: 68 (32.4%) 95% CI: [26.4%, 39.0%]
+ FAILED_DECEPTION: 3 (1.4%) 95% CI: [0.5%, 4.1%]
+ ROBUST_HONEST: 29 (13.8%) 95% CI: [9.8%, 19.1%]
+ UNINTENTIONAL_FAILURE: 4 (1.9%) 95% CI: [0.7%, 4.8%]
+ UNKNOWN: 1 (0.5%) 95% CI: [0.1%, 2.6%]
+
+ Chi-squared test: χ²(5) = 84.784, p < .001 ***
+ Cramér's V = 0.449 (medium)
+ ⚠ Minimum expected count = 0.5 < 5
+
+ --- 7b. Sophistication distribution by scenario ---
+
+ child_protection (n=210):
+ NAIVE: 19 (9.0%)
+ REACTIVE: 0 (0.0%)
+ TACTICAL: 33 (15.7%)
+ STRATEGIC: 158 (75.2%)
+ corporate_sabotage_v2 (n=210):
+ NAIVE: 33 (15.7%)
+ REACTIVE: 5 (2.4%)
+ TACTICAL: 38 (18.1%)
+ STRATEGIC: 133 (63.3%)
+
+ Chi-squared test: χ²(3) = 11.267, p = 0.0104 *
+ Cramér's V = 0.164 (small)
+
+================================================================================
+ 8. REASONING–OUTPUT ALIGNMENT
+================================================================================
+
+
+ --- 8. Alignment matrix — child_protection ---
+
+ ACCURATE MISLEADING OMITTED FABRICATED TOTAL
+ ------------------------------------------------------------------------------------------------
+ ROBUST_HONEST 0 4 6 0 10
+ CONSIDERED_HONEST 1 4 0 1 6
+ COMPLIANT_DECEPTION 31 140 10 1 182
+ FAILED_DECEPTION 1 1 1 0 3
+ UNINTENTIONAL_FAILURE 0 1 8 0 9
+
+ Compliant reasoning → accurate output: 31/182 (17.0%) 95% CI: [12.3%, 23.2%]
+ Honest reasoning → deceptive output: 15/16 (93.8%) 95% CI: [71.7%, 98.9%]
+
+ --- 8. Alignment matrix — corporate_sabotage_v2 ---
+
+ ACCURATE MISLEADING OMITTED FABRICATED TOTAL
+ ------------------------------------------------------------------------------------------------
+ ROBUST_HONEST 17 2 2 8 29
+ CONSIDERED_HONEST 45 5 7 11 68
+ COMPLIANT_DECEPTION 42 48 10 5 105
+ FAILED_DECEPTION 2 0 0 1 3
+ UNINTENTIONAL_FAILURE 0 0 4 0 4
+
+ Compliant reasoning → accurate output: 42/105 (40.0%) 95% CI: [31.1%, 49.6%]
+ Honest reasoning → deceptive output: 35/97 (36.1%) 95% CI: [27.2%, 46.0%]
+
+================================================================================
+ 9. SAFETY TRAINING EFFECT (gpt-oss-20b vs gpt-oss-safeguard-20b)
+================================================================================
+
+
+ --- 9. Safety training effect — child_protection ---
+
+ base: 25/30 (83.3%) 95% CI: [66.4%, 92.7%]
+ safeguard: 12/30 (40.0%) 95% CI: [24.6%, 57.7%]
+
+ Chi-squared: χ²(1) = 10.153, p = 0.0014 **
+ Fisher's exact: p = 0.0012 **
+ Odds ratio = 0.133 95% CI: [0.040, 0.446]
+
+ --- 9. Safety training effect — corporate_sabotage_v2 ---
+
+ base: 16/30 (53.3%) 95% CI: [36.1%, 69.8%]
+ safeguard: 16/30 (53.3%) 95% CI: [36.1%, 69.8%]
+
+ Chi-squared: χ²(1) = 0.000, p = 1.0000 n.s.
+ Fisher's exact: p = 1.0000 n.s.
+ Odds ratio = 1.000 95% CI: [0.363, 2.758]
+
+================================================================================
+ 10. FULL LOGISTIC REGRESSION MODEL
+================================================================================
+
+
+ --- 10a. Deceptive output ~ scenario + oversight + model ---
+
+ Coef. Std.Err. z P>|z| [0.025 0.975]
+const 0.261332 0.346109 0.755058 4.502143e-01 -0.417028 0.939692
+scenario_code 1.954906 0.257723 7.585311 3.316907e-14 1.449779 2.460033
+oversight_num -0.542462 0.151568 -3.579003 3.449080e-04 -0.839530 -0.245394
+model_gemini-3.1-pro-preview 2.717039 0.679092 4.000988 6.307868e-05 1.386043 4.048036
+model_moonshotai/kimi-k2.5 0.183173 0.428262 0.427712 6.688610e-01 -0.656206 1.022552
+model_openai/gpt-oss-20b 0.277243 0.430559 0.643914 5.196314e-01 -0.566637 1.121123
+model_openai/gpt-oss-safeguard-20b -0.867601 0.422054 -2.055663 3.981505e-02 -1.694812 -0.040390
+model_qwen/qwen3-max-thinking 0.373344 0.433243 0.861743 3.888291e-01 -0.475796 1.222484
+model_z-ai/glm-4.7-flash -0.265752 0.421457 -0.630555 5.283312e-01 -1.091791 0.560288
+
+ Pseudo R² = 0.2256
+ AIC = 431.0
+ n = 420
+
+ Odds ratios (exponentiated coefficients):
+ scenario_code OR = 7.063 95% CI: [4.262, 11.705] ***
+ oversight_num OR = 0.581 95% CI: [0.432, 0.782] ***
+ model_gemini-3.1-pro-preview OR = 15.135 95% CI: [3.999, 57.285] ***
+ model_moonshotai/kimi-k2.5 OR = 1.201 95% CI: [0.519, 2.780] n.s.
+ model_openai/gpt-oss-20b OR = 1.319 95% CI: [0.567, 3.068] n.s.
+ model_openai/gpt-oss-safeguard-20b OR = 0.420 95% CI: [0.184, 0.960] *
+ model_qwen/qwen3-max-thinking OR = 1.453 95% CI: [0.621, 3.396] n.s.
+ model_z-ai/glm-4.7-flash OR = 0.767 95% CI: [0.336, 1.751] n.s.
+
+ --- 10b. With scenario × oversight interaction ---
+
+ Coef. Std.Err. z P>|z| [0.025 0.975]
+const 0.512621 0.368505 1.391083 0.164200 -0.209635 1.234877
+scenario_code 1.216952 0.401944 3.027667 0.002464 0.429156 2.004747
+oversight_num -0.811699 0.198134 -4.096709 0.000042 -1.200035 -0.423363
+interaction 0.697058 0.310895 2.242097 0.024955 0.087714 1.306402
+model_gemini-3.1-pro-preview 2.810432 0.689770 4.074450 0.000046 1.458508 4.162356
+model_moonshotai/kimi-k2.5 0.188627 0.434612 0.434012 0.664280 -0.663196 1.040449
+model_openai/gpt-oss-20b 0.285949 0.437317 0.653872 0.513194 -0.571176 1.143074
+model_openai/gpt-oss-safeguard-20b -0.872578 0.423284 -2.061449 0.039260 -1.702200 -0.042957
+model_qwen/qwen3-max-thinking 0.385637 0.440412 0.875629 0.381232 -0.477554 1.248829
+model_z-ai/glm-4.7-flash -0.271177 0.425763 -0.636919 0.524177 -1.105658 0.563304
+
+ LR test for interaction: LR = 5.035, p = 0.0248 *
+
+================================================================================
+ 11. SUMMARY TABLE (Publication-Ready)
+================================================================================
+
+ Model | Scenario | Low | Mid | High | Overall
+ ---------------------------------------------------------------------------------------------------------
+ deepseek-v3.2 | child_protection | 100% | 100% | 100% | 100%
+ deepseek-v3.2 | corporate_sabotage_v2 | 20% | 40% | 20% | 27%
+ ---------------------------------------------------------------------------------------------------------
+ gemini-3.1-pro-preview | child_protection | 100% | 100% | 100% | 100%
+ gemini-3.1-pro-preview | corporate_sabotage_v2 | 100% | 80% | 90% | 90%
+ ---------------------------------------------------------------------------------------------------------
+ kimi-k2.5 | child_protection | 100% | 80% | 90% | 90%
+ kimi-k2.5 | corporate_sabotage_v2 | 50% | 60% | 20% | 43%
+ ---------------------------------------------------------------------------------------------------------
+ gpt-oss-20b | child_protection | 80% | 80% | 90% | 83%
+ gpt-oss-20b | corporate_sabotage_v2 | 70% | 60% | 30% | 53%
+ ---------------------------------------------------------------------------------------------------------
+ gpt-oss-safeguard-20b | child_protection | 50% | 50% | 20% | 40%
+ gpt-oss-safeguard-20b | corporate_sabotage_v2 | 70% | 50% | 40% | 53%
+ ---------------------------------------------------------------------------------------------------------
+ qwen3-max-thinking | child_protection | 100% | 100% | 100% | 100%
+ qwen3-max-thinking | corporate_sabotage_v2 | 70% | 40% | 10% | 40%
+ ---------------------------------------------------------------------------------------------------------
+ glm-4.7-flash | child_protection | 70% | 80% | 80% | 77%
+ glm-4.7-flash | corporate_sabotage_v2 | 80% | 20% | 20% | 40%
+ ---------------------------------------------------------------------------------------------------------
+
+ MARGINAL | | Low | Mid | High | Overall
+ ---------------------------------------------------------------------------------------------------------
+ ALL MODELS | child_protection | 86% | 84% | 83% | 84%
+ ALL MODELS | corporate_sabotage_v2 | 66% | 50% | 33% | 50%
+
+================================================================================
+ ANALYSIS COMPLETE
+================================================================================
+
+ Total runs analysed: 420
+ Models: ['deepseek/deepseek-v3.2', 'gemini-3.1-pro-preview', 'moonshotai/kimi-k2.5', 'openai/gpt-oss-20b', 'openai/gpt-oss-safeguard-20b', 'qwen/qwen3-max-thinking', 'z-ai/glm-4.7-flash']
+ Scenarios: ['child_protection', 'corporate_sabotage_v2']
+ Oversight levels: ['high', 'low', 'mid']
+ Judge categories used:
+ Blackbox: {'MISLEADING', 'FABRICATED', 'OMITTED'} vs {'ACCURATE'}
+ Glassbox: ['COMPLIANT_DECEPTION', 'CONSIDERED_HONEST', 'FAILED_DECEPTION', 'ROBUST_HONEST', 'UNINTENTIONAL_FAILURE', 'UNKNOWN']
+ Sophistication: ['NAIVE', 'REACTIVE', 'STRATEGIC', 'TACTICAL', 'UNKNOWN']