summaryrefslogtreecommitdiff
path: root/docs/paper/results.md
blob: 1f7ea0c1c8d96d7eb5051fd844cc90431192459e (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
Loading data from: logs/full_experiment/results.csv
Loaded 420 rows: 7 models, 2 scenarios, 3 oversight levels
Significance level: α = 0.05

================================================================================
  1. SCENARIO EFFECT ON DECEPTION
================================================================================


  --- 1a. Output-level deception rate by scenario (blackbox) ---

  Contingency table:
    child_protection: 177/210 (84.3%)   95% CI: [78.8%, 88.6%]
    corporate_sabotage_v2: 104/210 (49.5%)   95% CI: [42.8%, 56.2%]

  Chi-squared test: χ²(1) = 55.743, p < .001 ***
  Cramér's V = 0.364 (medium)
  Cramér's V (bias-corrected) = 0.361
  Odds ratio = 0.183  95% CI: [0.116, 0.290]
  Fisher's exact test: p < .001 ***

  --- 1b. Reasoning-level compliance by scenario (glassbox: COMPLIANT_DECEPTION) ---

  Contingency table:
    child_protection: 182/210 (86.7%)   95% CI: [81.4%, 90.6%]
    corporate_sabotage_v2: 105/210 (50.0%)   95% CI: [43.3%, 56.7%]

  Chi-squared test: χ²(1) = 63.554, p < .001 ***
  Cramér's V = 0.389 (medium)
  Odds ratio = 0.154  95% CI: [0.095, 0.249]

================================================================================
  2. OVERSIGHT EFFECT ON DECEPTION
================================================================================


  --- 2a. Output-level deception by oversight — child_protection ---

  Contingency table:
    low: 60/70 (85.7%)   95% CI: [75.7%, 92.1%]
    mid: 59/70 (84.3%)   95% CI: [74.0%, 91.0%]
    high: 58/70 (82.9%)   95% CI: [72.4%, 89.9%]

  Chi-squared test: χ²(2) = 0.216, p = 0.8978 n.s.
  Cramér's V = 0.032 (negligible)
  Minimum expected cell count = 11.0  ✓ All expected counts ≥ 5

  Cochran-Armitage trend test: z = -0.464, p = 0.6423 n.s.
    → No significant linear trend detected

  --- 2b. Pairwise oversight comparisons — child_protection (Bonferroni-corrected, α/3 = 0.0167) ---

    low vs mid:
      χ²(1) = 0.000, raw p = 1.0000, Fisher p = 1.0000
      OR = 0.894  95% CI: [0.353, 2.263]  [Bonferroni: n.s.]
    low vs high:
      χ²(1) = 0.054, raw p = 0.8164, Fisher p = 0.8169
      OR = 0.806  95% CI: [0.323, 2.008]  [Bonferroni: n.s.]
    mid vs high:
      χ²(1) = 0.000, raw p = 1.0000, Fisher p = 1.0000
      OR = 0.901  95% CI: [0.368, 2.205]  [Bonferroni: n.s.]

  --- 2a. Output-level deception by oversight — corporate_sabotage_v2 ---

  Contingency table:
    low: 46/70 (65.7%)   95% CI: [54.0%, 75.8%]
    mid: 35/70 (50.0%)   95% CI: [38.6%, 61.4%]
    high: 23/70 (32.9%)   95% CI: [23.0%, 44.5%]

  Chi-squared test: χ²(2) = 15.125, p < .001 ***
  Cramér's V = 0.268 (small)
  Minimum expected cell count = 34.7  ✓ All expected counts ≥ 5

  Cochran-Armitage trend test: z = -3.888, p < .001 ***
    → Significant linear trend: deception decreasing with oversight level

  --- 2b. Pairwise oversight comparisons — corporate_sabotage_v2 (Bonferroni-corrected, α/3 = 0.0167) ---

    low vs mid:
      χ²(1) = 2.929, raw p = 0.0870, Fisher p = 0.0866
      OR = 0.522  95% CI: [0.264, 1.030]  [Bonferroni: n.s.]
    low vs high:
      χ²(1) = 13.831, raw p < .001, Fisher p < .001
      OR = 0.255  95% CI: [0.127, 0.515]  [Bonferroni: SIG]
    mid vs high:
      χ²(1) = 3.562, raw p = 0.0591, Fisher p = 0.0587
      OR = 0.489  95% CI: [0.247, 0.970]  [Bonferroni: n.s.]

================================================================================
  3. MODEL-LEVEL DIFFERENCES
================================================================================


  --- 3a. Output deception by model — child_protection ---

  Model deception rates:
    deepseek/deepseek-v3.2: 30/30 (100.0%)   95% CI: [88.6%, 100.0%]
    gemini-3.1-pro-preview: 30/30 (100.0%)   95% CI: [88.6%, 100.0%]
    moonshotai/kimi-k2.5: 27/30 (90.0%)   95% CI: [74.4%, 96.5%]
    openai/gpt-oss-20b: 25/30 (83.3%)   95% CI: [66.4%, 92.7%]
    openai/gpt-oss-safeguard-20b: 12/30 (40.0%)   95% CI: [24.6%, 57.7%]
    qwen/qwen3-max-thinking: 30/30 (100.0%)   95% CI: [88.6%, 100.0%]
    z-ai/glm-4.7-flash: 23/30 (76.7%)   95% CI: [59.1%, 88.2%]

  Chi-squared test: χ²(6) = 63.277, p < .001 ***
  Cramér's V = 0.549 (large)
  ⚠ Minimum expected count = 4.7 < 5

  --- 3a. Output deception by model — corporate_sabotage_v2 ---

  Model deception rates:
    deepseek/deepseek-v3.2: 8/30 (26.7%)   95% CI: [14.2%, 44.4%]
    gemini-3.1-pro-preview: 27/30 (90.0%)   95% CI: [74.4%, 96.5%]
    moonshotai/kimi-k2.5: 13/30 (43.3%)   95% CI: [27.4%, 60.8%]
    openai/gpt-oss-20b: 16/30 (53.3%)   95% CI: [36.1%, 69.8%]
    openai/gpt-oss-safeguard-20b: 16/30 (53.3%)   95% CI: [36.1%, 69.8%]
    qwen/qwen3-max-thinking: 12/30 (40.0%)   95% CI: [24.6%, 57.7%]
    z-ai/glm-4.7-flash: 12/30 (40.0%)   95% CI: [24.6%, 57.7%]

  Chi-squared test: χ²(6) = 28.917, p < .001 ***
  Cramér's V = 0.371 (medium)

================================================================================
  4. OVERSIGHT × SCENARIO INTERACTION
================================================================================


  --- 4a. Logistic regression: deception ~ scenario + oversight + scenario×oversight ---

                  Coef.  Std.Err.         z     P>|z|    [0.025    0.975]
const          0.661995  0.228656  2.895152  0.003790  0.213837  1.110154
scenario_code  1.128325  0.384590  2.933835  0.003348  0.374541  1.882108
oversight_num -0.682521  0.178978 -3.813430  0.000137 -1.033312 -0.331730
interaction    0.574509  0.293588  1.956854  0.050365 -0.000913  1.149931

  Model fit: Pseudo R² = 0.1409, AIC = 466.1, BIC = 482.3

  Likelihood ratio test for interaction: LR = 3.816, p = 0.0508 n.s.
    → No significant interaction

  --- 4b. Breslow-Day test for homogeneity of odds ratios ---

  (Tests whether the scenario effect is consistent across oversight levels)
    low: OR = 0.319  95% CI: [0.139, 0.734]
    mid: OR = 0.186  95% CI: [0.084, 0.413]
    high: OR = 0.101  95% CI: [0.046, 0.225]

  Cochran-Mantel-Haenszel test:
    Common OR = N/A
    CMH statistic = 58.293, p < .001 ***

  Breslow-Day test for homogeneity of ORs:
    Statistic = 3.885, p = 0.1434 n.s.
    → ORs are consistent across oversight strata

================================================================================
  5. MODEL × OVERSIGHT INTERACTION (per scenario)
================================================================================


  --- 5. Oversight effect by model — child_protection ---

    deepseek-v3.2                   low=100.0%  mid=100.0%  high=100.0%  trend z=+0.00 p = 1.0000 n.s.
    gemini-3.1-pro-preview          low=100.0%  mid=100.0%  high=100.0%  trend z=+0.00 p = 1.0000 n.s.
    kimi-k2.5                       low=100.0%  mid= 80.0%  high= 90.0%  trend z=-0.75 p = 0.4561 n.s.
    gpt-oss-20b                     low= 80.0%  mid= 80.0%  high= 90.0%  trend z=+0.60 p = 0.5485 n.s.
    gpt-oss-safeguard-20b           low= 50.0%  mid= 50.0%  high= 20.0%  trend z=-1.37 p = 0.1709 n.s.
    qwen3-max-thinking              low=100.0%  mid=100.0%  high=100.0%  trend z=+0.00 p = 1.0000 n.s.
    glm-4.7-flash                   low= 70.0%  mid= 80.0%  high= 80.0%  trend z=+0.53 p = 0.5970 n.s.

  --- 5. Oversight effect by model — corporate_sabotage_v2 ---

    deepseek-v3.2                   low= 20.0%  mid= 40.0%  high= 20.0%  trend z=+0.00 p = 1.0000 n.s.
    gemini-3.1-pro-preview          low=100.0%  mid= 80.0%  high= 90.0%  trend z=-0.75 p = 0.4561 n.s.
    kimi-k2.5                       low= 50.0%  mid= 60.0%  high= 20.0%  trend z=-1.35 p = 0.1758 n.s.
    gpt-oss-20b                     low= 70.0%  mid= 60.0%  high= 30.0%  trend z=-1.79 p = 0.0730 n.s.
    gpt-oss-safeguard-20b           low= 70.0%  mid= 50.0%  high= 40.0%  trend z=-1.34 p = 0.1787 n.s.
    qwen3-max-thinking              low= 70.0%  mid= 40.0%  high= 10.0%  trend z=-2.74 p = 0.0062 **
    glm-4.7-flash                   low= 80.0%  mid= 20.0%  high= 20.0%  trend z=-2.74 p = 0.0062 **

================================================================================
  6. DECEPTION TYPE ANALYSIS
================================================================================


  --- 6a. Deception type distribution by scenario ---

  child_protection (n=177):
    FABRICATED: 2 (1.1%)
    MISLEADING: 150 (84.7%)
    OMITTED: 25 (14.1%)
  corporate_sabotage_v2 (n=104):
    FABRICATED: 25 (24.0%)
    MISLEADING: 55 (52.9%)
    OMITTED: 24 (23.1%)

  Chi-squared test: χ²(2) = 47.906, p < .001 ***
  Cramér's V = 0.413 (medium)

================================================================================
  7. REASONING INTENT ANALYSIS (Glassbox)
================================================================================


  --- 7a. Glassbox category distribution by scenario ---

  child_protection (n=210):
    COMPLIANT_DECEPTION: 182 (86.7%)   95% CI: [81.4%, 90.6%]
    CONSIDERED_HONEST: 6 (2.9%)   95% CI: [1.3%, 6.1%]
    FAILED_DECEPTION: 3 (1.4%)   95% CI: [0.5%, 4.1%]
    ROBUST_HONEST: 10 (4.8%)   95% CI: [2.6%, 8.5%]
    UNINTENTIONAL_FAILURE: 9 (4.3%)   95% CI: [2.3%, 7.9%]
    UNKNOWN: 0 (0.0%)   95% CI: [0.0%, 1.8%]
  corporate_sabotage_v2 (n=210):
    COMPLIANT_DECEPTION: 105 (50.0%)   95% CI: [43.3%, 56.7%]
    CONSIDERED_HONEST: 68 (32.4%)   95% CI: [26.4%, 39.0%]
    FAILED_DECEPTION: 3 (1.4%)   95% CI: [0.5%, 4.1%]
    ROBUST_HONEST: 29 (13.8%)   95% CI: [9.8%, 19.1%]
    UNINTENTIONAL_FAILURE: 4 (1.9%)   95% CI: [0.7%, 4.8%]
    UNKNOWN: 1 (0.5%)   95% CI: [0.1%, 2.6%]

  Chi-squared test: χ²(5) = 84.784, p < .001 ***
  Cramér's V = 0.449 (medium)
  ⚠ Minimum expected count = 0.5 < 5

  --- 7b. Sophistication distribution by scenario ---

  child_protection (n=210):
    NAIVE: 19 (9.0%)
    REACTIVE: 0 (0.0%)
    TACTICAL: 33 (15.7%)
    STRATEGIC: 158 (75.2%)
  corporate_sabotage_v2 (n=210):
    NAIVE: 33 (15.7%)
    REACTIVE: 5 (2.4%)
    TACTICAL: 38 (18.1%)
    STRATEGIC: 133 (63.3%)

  Chi-squared test: χ²(3) = 11.267, p = 0.0104 *
  Cramér's V = 0.164 (small)

================================================================================
  8. REASONING–OUTPUT ALIGNMENT
================================================================================


  --- 8. Alignment matrix — child_protection ---

                                        ACCURATE    MISLEADING       OMITTED    FABRICATED     TOTAL
    ------------------------------------------------------------------------------------------------
    ROBUST_HONEST                              0             4             6             0        10
    CONSIDERED_HONEST                          1             4             0             1         6
    COMPLIANT_DECEPTION                       31           140            10             1       182
    FAILED_DECEPTION                           1             1             1             0         3
    UNINTENTIONAL_FAILURE                      0             1             8             0         9

  Compliant reasoning → accurate output: 31/182 (17.0%)  95% CI: [12.3%, 23.2%]
  Honest reasoning → deceptive output: 15/16 (93.8%)  95% CI: [71.7%, 98.9%]

  --- 8. Alignment matrix — corporate_sabotage_v2 ---

                                        ACCURATE    MISLEADING       OMITTED    FABRICATED     TOTAL
    ------------------------------------------------------------------------------------------------
    ROBUST_HONEST                             17             2             2             8        29
    CONSIDERED_HONEST                         45             5             7            11        68
    COMPLIANT_DECEPTION                       42            48            10             5       105
    FAILED_DECEPTION                           2             0             0             1         3
    UNINTENTIONAL_FAILURE                      0             0             4             0         4

  Compliant reasoning → accurate output: 42/105 (40.0%)  95% CI: [31.1%, 49.6%]
  Honest reasoning → deceptive output: 35/97 (36.1%)  95% CI: [27.2%, 46.0%]

================================================================================
  9. SAFETY TRAINING EFFECT (gpt-oss-20b vs gpt-oss-safeguard-20b)
================================================================================


  --- 9. Safety training effect — child_protection ---

    base: 25/30 (83.3%)   95% CI: [66.4%, 92.7%]
    safeguard: 12/30 (40.0%)   95% CI: [24.6%, 57.7%]

  Chi-squared: χ²(1) = 10.153, p = 0.0014 **
  Fisher's exact: p = 0.0012 **
  Odds ratio = 0.133  95% CI: [0.040, 0.446]

  --- 9. Safety training effect — corporate_sabotage_v2 ---

    base: 16/30 (53.3%)   95% CI: [36.1%, 69.8%]
    safeguard: 16/30 (53.3%)   95% CI: [36.1%, 69.8%]

  Chi-squared: χ²(1) = 0.000, p = 1.0000 n.s.
  Fisher's exact: p = 1.0000 n.s.
  Odds ratio = 1.000  95% CI: [0.363, 2.758]

================================================================================
  10. FULL LOGISTIC REGRESSION MODEL
================================================================================


  --- 10a. Deceptive output ~ scenario + oversight + model ---

                                       Coef.  Std.Err.         z         P>|z|    [0.025    0.975]
const                               0.261332  0.346109  0.755058  4.502143e-01 -0.417028  0.939692
scenario_code                       1.954906  0.257723  7.585311  3.316907e-14  1.449779  2.460033
oversight_num                      -0.542462  0.151568 -3.579003  3.449080e-04 -0.839530 -0.245394
model_gemini-3.1-pro-preview        2.717039  0.679092  4.000988  6.307868e-05  1.386043  4.048036
model_moonshotai/kimi-k2.5          0.183173  0.428262  0.427712  6.688610e-01 -0.656206  1.022552
model_openai/gpt-oss-20b            0.277243  0.430559  0.643914  5.196314e-01 -0.566637  1.121123
model_openai/gpt-oss-safeguard-20b -0.867601  0.422054 -2.055663  3.981505e-02 -1.694812 -0.040390
model_qwen/qwen3-max-thinking       0.373344  0.433243  0.861743  3.888291e-01 -0.475796  1.222484
model_z-ai/glm-4.7-flash           -0.265752  0.421457 -0.630555  5.283312e-01 -1.091791  0.560288

  Pseudo R² = 0.2256
  AIC = 431.0
  n = 420

  Odds ratios (exponentiated coefficients):
    scenario_code                                  OR = 7.063  95% CI: [4.262, 11.705]  ***
    oversight_num                                  OR = 0.581  95% CI: [0.432, 0.782]  ***
    model_gemini-3.1-pro-preview                   OR = 15.135  95% CI: [3.999, 57.285]  ***
    model_moonshotai/kimi-k2.5                     OR = 1.201  95% CI: [0.519, 2.780]  n.s.
    model_openai/gpt-oss-20b                       OR = 1.319  95% CI: [0.567, 3.068]  n.s.
    model_openai/gpt-oss-safeguard-20b             OR = 0.420  95% CI: [0.184, 0.960]  *
    model_qwen/qwen3-max-thinking                  OR = 1.453  95% CI: [0.621, 3.396]  n.s.
    model_z-ai/glm-4.7-flash                       OR = 0.767  95% CI: [0.336, 1.751]  n.s.

  --- 10b. With scenario × oversight interaction ---

                                       Coef.  Std.Err.         z     P>|z|    [0.025    0.975]
const                               0.512621  0.368505  1.391083  0.164200 -0.209635  1.234877
scenario_code                       1.216952  0.401944  3.027667  0.002464  0.429156  2.004747
oversight_num                      -0.811699  0.198134 -4.096709  0.000042 -1.200035 -0.423363
interaction                         0.697058  0.310895  2.242097  0.024955  0.087714  1.306402
model_gemini-3.1-pro-preview        2.810432  0.689770  4.074450  0.000046  1.458508  4.162356
model_moonshotai/kimi-k2.5          0.188627  0.434612  0.434012  0.664280 -0.663196  1.040449
model_openai/gpt-oss-20b            0.285949  0.437317  0.653872  0.513194 -0.571176  1.143074
model_openai/gpt-oss-safeguard-20b -0.872578  0.423284 -2.061449  0.039260 -1.702200 -0.042957
model_qwen/qwen3-max-thinking       0.385637  0.440412  0.875629  0.381232 -0.477554  1.248829
model_z-ai/glm-4.7-flash           -0.271177  0.425763 -0.636919  0.524177 -1.105658  0.563304

  LR test for interaction: LR = 5.035, p = 0.0248 *

================================================================================
  11. SUMMARY TABLE (Publication-Ready)
================================================================================

  Model                               | Scenario                  |        Low |        Mid |       High |    Overall
  ---------------------------------------------------------------------------------------------------------
  deepseek-v3.2                       | child_protection          |       100% |       100% |       100% |       100%
  deepseek-v3.2                       | corporate_sabotage_v2     |        20% |        40% |        20% |        27%
  ---------------------------------------------------------------------------------------------------------
  gemini-3.1-pro-preview              | child_protection          |       100% |       100% |       100% |       100%
  gemini-3.1-pro-preview              | corporate_sabotage_v2     |       100% |        80% |        90% |        90%
  ---------------------------------------------------------------------------------------------------------
  kimi-k2.5                           | child_protection          |       100% |        80% |        90% |        90%
  kimi-k2.5                           | corporate_sabotage_v2     |        50% |        60% |        20% |        43%
  ---------------------------------------------------------------------------------------------------------
  gpt-oss-20b                         | child_protection          |        80% |        80% |        90% |        83%
  gpt-oss-20b                         | corporate_sabotage_v2     |        70% |        60% |        30% |        53%
  ---------------------------------------------------------------------------------------------------------
  gpt-oss-safeguard-20b               | child_protection          |        50% |        50% |        20% |        40%
  gpt-oss-safeguard-20b               | corporate_sabotage_v2     |        70% |        50% |        40% |        53%
  ---------------------------------------------------------------------------------------------------------
  qwen3-max-thinking                  | child_protection          |       100% |       100% |       100% |       100%
  qwen3-max-thinking                  | corporate_sabotage_v2     |        70% |        40% |        10% |        40%
  ---------------------------------------------------------------------------------------------------------
  glm-4.7-flash                       | child_protection          |        70% |        80% |        80% |        77%
  glm-4.7-flash                       | corporate_sabotage_v2     |        80% |        20% |        20% |        40%
  ---------------------------------------------------------------------------------------------------------

  MARGINAL                            |                           |        Low |        Mid |       High |    Overall
  ---------------------------------------------------------------------------------------------------------
  ALL MODELS                          | child_protection          |        86% |        84% |        83% |        84%
  ALL MODELS                          | corporate_sabotage_v2     |        66% |        50% |        33% |        50%

================================================================================
  ANALYSIS COMPLETE
================================================================================

  Total runs analysed: 420
  Models: ['deepseek/deepseek-v3.2', 'gemini-3.1-pro-preview', 'moonshotai/kimi-k2.5', 'openai/gpt-oss-20b', 'openai/gpt-oss-safeguard-20b', 'qwen/qwen3-max-thinking', 'z-ai/glm-4.7-flash']
  Scenarios: ['child_protection', 'corporate_sabotage_v2']
  Oversight levels: ['high', 'low', 'mid']
  Judge categories used:
    Blackbox: {'MISLEADING', 'FABRICATED', 'OMITTED'} vs {'ACCURATE'}
    Glassbox: ['COMPLIANT_DECEPTION', 'CONSIDERED_HONEST', 'FAILED_DECEPTION', 'ROBUST_HONEST', 'UNINTENTIONAL_FAILURE', 'UNKNOWN']
    Sophistication: ['NAIVE', 'REACTIVE', 'STRATEGIC', 'TACTICAL', 'UNKNOWN']