summaryrefslogtreecommitdiff
path: root/projects/dis
diff options
context:
space:
mode:
authorCaptainJack2491 <jayrupnakawala@gmail.com>2026-07-07 15:41:59 +0100
committerCaptainJack2491 <jayrupnakawala@gmail.com>2026-07-07 15:41:59 +0100
commit02c429e5ed53fe42a0a1f3025a16a0dc5518ad36 (patch)
tree0ac75d2be9c4cec332f6ba5802906b1c2a11146f /projects/dis
parent3bf5ad40c4ee63f9c2f53019fdbede712521bc32 (diff)
update CV with dissertation findings, add dissertation page, add website project entry
Diffstat (limited to 'projects/dis')
-rw-r--r--projects/dis/index.qmd79
1 files changed, 79 insertions, 0 deletions
diff --git a/projects/dis/index.qmd b/projects/dis/index.qmd
new file mode 100644
index 0000000..ef62874
--- /dev/null
+++ b/projects/dis/index.qmd
@@ -0,0 +1,79 @@
+---
+title: "Dissertation"
+---
+
+## MORAL COVER AND WATCHFUL EYES: THE LIMITS OF IN-CONTEXT INTERVENTIONS ON LLM DECEPTION
+
+Jayrup Nakawala \
+BSc Data Science and Artificial Intelligence \
+University of East London — May 2026 \
+Supervisor: Dr. Aloysius Adotey Edoh
+
+--------------------------------------------------
+
+## ABSTRACT
+
+As LLMs are entrusted with real-world tool use, the
+question of whether they can be reliably monitored
+for deception becomes a practical safety concern.
+
+This dissertation studies whether in-context oversight
+cues and hidden-goal framing can meaningfully affect
+deceptive behaviour in agentic LLMs. A sandboxed
+evaluation framework lets models interact with a
+virtual file system and assesses them using three
+complementary signals:
+
+- deterministic regex checks
+- blackbox output judgment
+- glassbox reasoning-trace judgment
+
+960 runs across three stages: a 420-run pilot over
+seven models, followed by two 270-run confirmatory
+studies over three models.
+
+--------------------------------------------------
+
+## KEY FINDINGS
+
+1. Oversight does not deter deception.
+ Increasing perceived oversight had no significant
+ effect on deception frequency (p = 0.587).
+ Models process oversight intensively in their
+ reasoning but it produces no behavioural change.
+
+2. Moral framing does not increase deception.
+ Rates were identical across bare (66.7%), moral
+ (65.6%), and self-serving (66.7%) framings.
+
+3. The Stealth Gap is real and structural.
+ Output-based detection underestimates deceptive
+ intent by 35–46% at every oversight level.
+ Models pass oversight evaluations through superior
+ concealment, not genuine alignment.
+
+4. Model identity dominates everything.
+ Model choice explained 60–83% of variance in
+ deception. Some models deceive at near-ceiling
+ rates; others refuse entirely — regardless of
+ experimental condition.
+
+--------------------------------------------------
+
+## RESEARCH QUESTIONS
+
+RQ1: How does perceived oversight affect agentic
+ deception? \
+RQ2: Does moral framing of conflicting objectives
+ influence deceptive behaviour? \
+RQ3: Does oversight reduce intent or just drive
+ concealment? \
+RQ4: Can multi-pronged evaluation (regex + blackbox
+ + glassbox) provide robust classifications?
+
+--------------------------------------------------
+
+DOWNLOAD: curl [jayrup.me/dissertation.pdf](https://jayrup.me/dissertation.pdf) \
+CODE: [github.com/CaptainJack2491/dissertation](https://github.com/CaptainJack2491/dissertation)
+
+--------------------------------------------------