summaryrefslogtreecommitdiff
path: root/notes
diff options
context:
space:
mode:
authorCaptainJack2491 <jayrupnakawala@gmail.com>2025-11-16 18:22:05 +0000
committerCaptainJack2491 <jayrupnakawala@gmail.com>2025-11-16 18:22:05 +0000
commit0a24528267e7e67be5d62daefbd061a98670eb15 (patch)
tree4db0da029ba2bb78b318b5c3f464d730a6a451d4 /notes
parentcb1012d7464bb5b3a319fb4ea85dbb1178861ddb (diff)
Move docs and notes and papers from feature-branch to main
Diffstat (limited to 'notes')
-rw-r--r--notes/approach_documentation.md31
-rw-r--r--notes/b.md11
2 files changed, 42 insertions, 0 deletions
diff --git a/notes/approach_documentation.md b/notes/approach_documentation.md
new file mode 100644
index 0000000..077d20e
--- /dev/null
+++ b/notes/approach_documentation.md
@@ -0,0 +1,31 @@
+# Capturing Pre-Tool-Call Reasoning from Language Models
+
+## Problem
+
+The goal is to capture the reasoning process of a language model *before* it decides to call a tool. This "pre-tool-call reasoning" is crucial for understanding the model's decision-making process, especially in the context of research on agentic behavior and alignment. High-level library abstractions for tool calling often hide this part of the model's output, focusing only on the tool call itself.
+
+## Chosen Approach: Direct Model Invocation with LangChain
+
+To address this, we have adopted a lower-level approach within the `langchain` ecosystem. Instead of using high-level abstractions like `bind_tools`, we interact with the `ChatOllama` model more directly. This approach gives us the necessary control to access the raw output from the model and parse it according to our specific needs.
+
+## Implementation Details
+
+The implementation in `src/agents/02-sandbox/main.py` follows these steps:
+
+1. **Manual Prompt Construction**: We create a detailed system prompt that explicitly instructs the model to first "think" about the problem and write down its reasoning in a `<think>` block, and then to output the tool call in a `<tool_call>` block. The tool definitions are rendered as text and included in the prompt.
+
+2. **Direct Model Invocation**: We use the `llm.invoke()` method to send the prompt to the model and receive the raw `AIMessage` response. This response contains the model's output as a single string, including our custom `<think>` and `<tool_call>` blocks.
+
+3. **Response Parsing**: The script then parses this raw response using regular expressions to extract the content of the `<think>` and `<tool_call>` blocks separately.
+
+4. **Tool Execution**: After parsing the tool call, the script identifies the corresponding tool function and executes it with the provided arguments.
+
+## Rationale
+
+This approach was chosen for the following reasons:
+
+- **Control and Transparency**: It provides full control over the model's output, allowing us to capture the valuable reasoning tokens that are often lost when using high-level abstractions.
+- **Ecosystem Alignment**: It stays within the `langchain` ecosystem, which is already in use for the project. This allows us to leverage `langchain`'s strengths, such as multi-provider support and integration with logging and tracing tools like `LangSmith`, which are essential for the research project.
+- **Flexibility**: This method is highly flexible and can be adapted to different models and output formats. The parsing logic can be encapsulated into a custom `langchain` `BaseOutputParser` for better code organization and reusability.
+
+This approach successfully addresses the challenge of capturing pre-tool-call reasoning and provides a solid foundation for the experimental work in the project.
diff --git a/notes/b.md b/notes/b.md
new file mode 100644
index 0000000..71e2123
--- /dev/null
+++ b/notes/b.md
@@ -0,0 +1,11 @@
+- [ ] Agents and its connections to the main llms and tools/memory
+- [ ] parsing and logging and execution of tools
+
+
+
+- [ ] abstract
+- [ ] intro
+- [ ] Lit review (background, research gap)
+- [ ] methodology
+- [ ] implementation and results
+- [ ] conclusion/discussion (future works)