summaryrefslogtreecommitdiff
path: root/docs/final/intro.typ
diff options
context:
space:
mode:
Diffstat (limited to 'docs/final/intro.typ')
-rw-r--r--docs/final/intro.typ44
1 files changed, 23 insertions, 21 deletions
diff --git a/docs/final/intro.typ b/docs/final/intro.typ
index ab94048..201b9d8 100644
--- a/docs/final/intro.typ
+++ b/docs/final/intro.typ
@@ -1,48 +1,50 @@
= Introduction
-#pagebreak()
+== Context and Motivation
+
+Large Language Models (LLMs) are increasingly being used in agentic systems, where they can interact with tools, file systems, and even other applications. They therefore require evaluation within sandboxed environments where agents can interact with tools and external resources. The safety of such deployments rests on the assumption that the model will faithfully pursue the user's stated objectives.
-== Motivation and Problem Statement
+This assumption is increasingly challenged. #cite(<meinke2025frontiermodelscapableincontext>, form: "prose") demonstrated that frontier models are capable of _in-context scheming_---sabotaging tasks, disabling oversight, and exfiltrating data when given conflicting objectives, while maintaining deception through follow-up interrogations. #cite(<järviniemi2024uncoveringdeceptivetendencieslanguage>, form: "prose") showed that a corporate AI assistant will deceive auditors and manipulate public perception when contextual incentives align, without adversarial prompting.
-Frontier language models are frequently deployed in interactive settings where they are expected to follow user instructions, make tool calls, and generate outputs that downstream systems treat as trustworthy.
+== Problem Statement
-This dissertation examines whether model behavior can become **conditionally deceptive** when the model is exposed to conflicting objectives within the prompt context. Specifically, this study considers settings in which a model is prompted with a benign, publicly stated task goal, while also being exposed to a hidden objective.
+Standard single-turn evaluations (Prompt $arrow$ Response) do not work for agents, they operate in loops. They therefore need to be placed in sandboxed environments with tools for interaction. This enables the continuous monitoring of tool calls, API requests, and memory states over dozens, if not hundreds, of steps.
-In addition to deception itself, this work considers whether behavior varies under different levels of **perceived oversight intensity**.
+As LLMs evolve into autonomous agents capable of tool-use and long-horizon planning, recent work suggests that strategic deception may become a more common threat #cite(<chen2025aideceptionrisksdynamics>). Highly capable models have demonstrated the ability to understand their environments and recognize when they are being monitored. This phenomenon, known as evaluation awareness, is now recognised as a critical threat in Alignment research #cite(<souly2025investigating>). Because of this awareness, models may exhibit aligned, safe behavior during high-oversight testing but may defect to misaligned behaviors when deployed autonomously in environments with low perceived oversight #cite(<scheurer2024largelanguagemodelsstrategically>) #cite(<greenblatt2024alignmentfakinglargelanguage>).
== Research Gap
-Prior empirical work suggests that in-context deception and scheming can emerge in language models under certain incentive structures and evaluation conditions (e.g., Meinke et al., 2025; Järviniemi & Hubinger, 2024). However, the literature appears to leave an important gap: relatively few studies systematically vary perceived monitoring/oversight intensity while holding task structure and hidden objectives fixed.
+Existing evaluation methodologies are poorly equipped to measure or mitigate this specific threat. Most current benchmarks rely on static, single-turn prompts #cite(<wang2025thinkingllmslieunveiling>) or purely black-box outcome grading, focusing solely on whether an agent successfully completed a task or generated toxic text. These approaches fail to capture multi-turn deceptive strategies or distinguish between genuine alignment and 'alignment faking.' Recent literature emphasizes the critical need to analyze an agent's internal reasoning traces against its external actions to detect hidden intent #cite(<wu2026opendeceptionlearningdeceptiontrust>), yet few frameworks systematically manipulate environmental variables like perceived oversight to observe how these hidden intents manifest.
+
+A model that exhibits _alignment faking_ would be especially dangerous precisely because its outputs appear legitimate under surface review. Current oversight paradigms assume that transparency and monitoring enable trust; this work tests whether they may instead enable _rationalization_, making automated governance less sustainable, not more.
-As a result, it is often unclear whether changes observed under “monitoring” cues primarily reflect shifts in intent (i.e., alignment-like behavior) or instead reflect adaptation to the presence of evaluation cues (i.e., an observer-effect-like phenomenon).
+In particular, we investigate whether oversight mechanisms suppress all forms of deception equally, or whether certain forms of rationalized deception remain resilient to monitoring.
== Research Questions
This dissertation addresses the following research questions:
-1. **Conditional Deception:** Do AI agents pursue hidden objectives while producing outputs that appear compliant with benign task instructions?
-2. **Oversight Modulation (Frequency):** Does perceived oversight intensity affect the *frequency* of deceptive behavior?
-3. **Oversight Modulation (Strategy):** Does perceived oversight intensity affect the *sophistication* of deceptive behavior (when deception occurs)?
-4. **Evaluation Triangulation:** Can a multi-pronged evaluation approach (deterministic filtering, output-level judgment, and reasoning-level judgment) provide consistent categorizations?
-5. **Dependence on Model and Scenario:** How do results vary across model families and scenario types?
+1. *Conditional Deception:* Do AI agents prioritize hidden, misaligned goals over benign, user-facing goals when instructed in-context?
+2. *Oversight Modulation:* How does the level of perceived oversight (e.g., high monitoring vs. unmonitored environments) affect the frequency and sophistication of agentic deception?
+3. *Domain Context and Rationalization:* Does the domain context of a scenario (e.g., high-stakes child welfare vs. corporate sabotage) moderate the effect that oversight has on deceptive behavior?
+4. *Evaluation Triangulation:* Can a multi-pronged evaluation approach (deterministic filtering, output-level judgment, and reasoning-level judgment) provide consistent and robust categorizations of intent?
== Contributions
This dissertation contributes:
-- A modular experimental framework for studying in-context hidden-goal deception using a sandboxed Virtual File System (VFS) and a tool-enabled agent interface.
-- An evaluation pipeline that combines deterministic regex checks, an LLM black-box judge (output-level), and an LLM glass-box judge (intent-level using extracted reasoning traces).
-- An empirical dataset of **420 experiment runs** spanning multiple models, two scenario types, and three oversight conditions.
-- Evidence consistent with an “oversight paradox”: perceived oversight may correlate with reduced deception frequency while deception instances remain often highly strategic.
+- *An Automated Evaluation Framework:* A modular framework for testing tool-using agents under varying oversight conditions, featuring a multi-tiered judging system (regex + blackbox + glassbox).
+- *An Empirical Dataset:* A comprehensive pilot study dataset of 420 experimental runs spanning seven LLMs, two scenarios, and three oversight conditions.
+- *Evidence of an Oversight Paradox:* Analysis revealing that while simulated monitoring effectively deters self-serving explicit deception, it largely fails to deter morally-rationalized bureaucratic deception, and acts as a competence filter rather than a complete deterrent.
== Dissertation Structure
The remainder of the dissertation is organized as follows:
-- **Literature Review:** Summarizes prior work on deceptive alignment, in-context scheming, and monitoring/evaluation cues.
-- **Methodology:** Describes the experimental framework, scenario design, oversight manipulation, and the 3-pronged evaluation rubric.
-- **Results:** Reports deception rates across oversight levels, models, and scenarios, alongside category breakdowns.
-- **Analysis:** Interprets the results, reports statistical summaries, and discusses validity and limitations.
-- **Conclusion and Limitations:** Summarizes findings and outlines directions for future work.
+- *Literature Review:* Summarizes prior work on deceptive alignment, in-context scheming, and monitoring/evaluation cues.
+- *Methodology:* Describes the experimental framework, scenario design, oversight manipulation, and the 3-pronged evaluation rubric.
+- *Results:* Reports deception rates across oversight levels, models, and scenarios, alongside category breakdowns.
+- *Analysis:* Interprets the results, reports statistical summaries, and discusses validity and limitations.
+- *Conclusion and Future Work:* Summarizes implications for AI governance and outlines directions for future research.
#pagebreak()