summaryrefslogtreecommitdiff
path: root/notes/v2_redesign/07_judge_validation_results.md
diff options
context:
space:
mode:
authorCaptainJack2491 <jayrupnakawala@gmail.com>2026-04-15 21:18:57 +0100
committerCaptainJack2491 <jayrupnakawala@gmail.com>2026-04-15 21:18:57 +0100
commit6cc92304379996f68bb2df96a5e4b637b4c4804b (patch)
tree6cdbc8b5484934927cd6ac989e91fef44a60bb4c /notes/v2_redesign/07_judge_validation_results.md
parentd25f8fc2e2b02c04489ca1960b83ff7a70e7317a (diff)
Study 1: 270 runs complete + judge validation pipeline
- config_study1.yaml: 3 models × 3 oversight × bare framing × n=30 - Judge validation: 54-run subset, gold=Sonnet 4.6, proxy=Grok 4.1 Fast (BB κ=0.702) + GPT-4.1 (GB κ=0.878) - Fix OpenAI batch provider: BytesIO, method/url fields, response.body parsing - Scripts: extract_subset.py, judge_validation.py - Dissertation chapters updated (intro, methodology, results, conclusion)
Diffstat (limited to 'notes/v2_redesign/07_judge_validation_results.md')
-rw-r--r--notes/v2_redesign/07_judge_validation_results.md52
1 files changed, 52 insertions, 0 deletions
diff --git a/notes/v2_redesign/07_judge_validation_results.md b/notes/v2_redesign/07_judge_validation_results.md
new file mode 100644
index 0000000..faf32de
--- /dev/null
+++ b/notes/v2_redesign/07_judge_validation_results.md
@@ -0,0 +1,52 @@
+# Judge Validation Results — 2026-04-15
+
+## Study 1 Subset
+
+- 54 runs extracted from Study 1 (20% stratified subset out of 270, seed=42)
+- Logs: `logs/v2_study1_judge_subset/`
+
+## Gold Standard
+
+- Blackbox: `claude-sonnet-4-6` (Anthropic)
+- Glassbox: `claude-sonnet-4-6` (Anthropic)
+
+## Results
+
+| Model | BB κ | GB κ | Sph κ | Pass |
+| --------------------------- | --------- | --------- | --------- | ------- |
+| gpt-4.1 | 0.351 | 0.878 | 0.538 | no |
+| gpt-4.1-mini | 0.202 | 0.834 | 0.345 | no |
+| **grok-4-1-fast-reasoning** | **0.702** | **0.812** | **0.108** | **YES** |
+| grok-4.20-reasoning | 0.549 | 0.919 | 0.198 | no |
+| grok-4.20-non-reasoning | 0.390 | 0.878 | 0.117 | no |
+
+## Selected Proxy Judges
+
+Split configuration — best model per prong:
+
+| Prong | Proxy Judge | κ vs Gold | Rationale |
+| ------------ | ----------------------------- | ---------------------------- | ------------------------------- |
+| **Blackbox** | grok-4-1-fast-reasoning (xAI) | 0.702 | Only model to pass BB threshold |
+| **Glassbox** | gpt-4.1 (OpenAI) | 0.878 (intent), 0.538 (soph) | Best sophistication agreement |
+
+### Config
+
+```yaml
+judge:
+ blackbox:
+ model: grok-4-1-fast-reasoning
+ provider: xai
+ temperature: 0
+ glassbox:
+ model: gpt-4.1
+ provider: openai
+ temperature: 0
+```
+
+## Notes
+
+- All models show high glassbox agreement (κ > 0.8) — intent classification is consistent across judges
+- Blackbox agreement is more variable — output framing/categorization is harder to align on
+- Sophistication agreement is low across the board — tier classification is noisy (flagged as limitation)
+- Three different provider families used: Anthropic (gold), xAI (BB proxy), OpenAI (GB proxy) — no same-family bias
+- Gold standard used a single model (Claude Sonnet 4.6) for both prongs for consistency