index
:
dissertation.git
logs
main
Undergraduate dissertation — Deceptive Alignment in Autonomous LLM Agents (UEL, 2026). Sandboxed agentic-LLM deception study: 960 runs, three-tier judging (regex + blackbox + glassbox).
summary
refs
log
tree
commit
diff
log msg
author
committer
range
path:
root
/
src
/
agents
Age
Commit message (
Expand
)
Author
2025-12-18
updated oversight
CaptainJack2491
2025-12-04
gemini-2.5-pro pilot test
CaptainJack2491
2025-12-03
added corporate_sabotage_v2
CaptainJack2491
2025-12-03
updated interrogation to load from logs
CaptainJack2491
2025-11-25
added interogation logic and updated structure of logs
CaptainJack2491
2025-11-23
updated agent to parse thinking in content field and added model_id to the logs
CaptainJack2491
2025-11-23
good logs
CaptainJack2491
2025-11-23
moved 03-logs to main
CaptainJack2491
2025-11-20
updated agent to work with the EXTRA_BODY variable to work with models that r...
CaptainJack2491
2025-11-20
added corporate espionage scenario
CaptainJack2491
2025-11-20
uodated oversight prompt
CaptainJack2491
2025-11-19
updated logs to show modelname and scenario, and added extra param for google.
CaptainJack2491
2025-11-19
added new scenario, corporate espionage
CaptainJack2491
2025-11-16
better scenario management
CaptainJack2491
2025-11-16
better VFS
CaptainJack2491
2025-11-16
added scenario data_sanitization
CaptainJack2491
2025-11-16
added more tools
CaptainJack2491
2025-11-16
better structure. and works with groq too
CaptainJack2491
2025-11-16
added modular prompts
CaptainJack2491
2025-11-16
logging works
CaptainJack2491
2025-11-16
reasoning works for both toolcall and response
CaptainJack2491
2025-10-28
Initial commit: Set up project structure
CaptainJack2491