Skip to content

Commit 7d86921

Browse files
femtoclaude
andcommitted
Add documentation for Auto-compact and Tool Search Tool (TST)
- docs/auto_compact.md: Automatic context window management via history summarization - docs/tool_search.md: Dynamic tool discovery for large tool libraries (85% token reduction) - README.md: Add links to new documentation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
1 parent 1d45cfa commit 7d86921

3 files changed

Lines changed: 483 additions & 0 deletions

File tree

‎README.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -139,6 +139,8 @@ The flowchart demonstrates the complete process from query to final result:
139139
- [Skills Guide](docs/skills.md) - Extend agent capabilities with modular skills
140140
- [Benchmarks](docs/benchmarks.md) - Performance results on GSM8K, Game of 24, AIME, Humaneval
141141
- [Route Parameter Guide](docs/agent_route_parameter_guide.md) - Route options for different reasoning strategies
142+
- [Tool Search Tool (TST)](docs/tool_search.md) - Dynamic tool discovery for large tool libraries (85% token reduction)
143+
- [Auto-compact Guide](docs/auto_compact.md) - Automatic context window management via history summarization
142144
- [Auto-decay Guide](docs/auto_decay.md) - Automatic context management for large tool responses (Experimental)
143145
144146
## Configuration

‎docs/auto_compact.md‎

Lines changed: 206 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,206 @@
1+
# Auto-compact: Automatic Context Window Management
2+
3+
## Overview
4+
5+
As agent conversations grow longer, the context window fills up with historical messages. **Auto-compact** automatically manages context by summarizing older messages when the conversation approaches the context limit.
6+
7+
**Auto-compact** helps prevent:
8+
- Context overflow errors
9+
- Performance degradation from processing excessive tokens
10+
- Increased API costs from large context windows
11+
12+
## How It Works
13+
14+
### 1. Monitor Phase
15+
16+
After each step, the system monitors context usage:
17+
- Calculates current token count
18+
- Compares against the model's context window limit
19+
- Triggers compaction when usage exceeds threshold (default: 92%)
20+
21+
### 2. Compact Phase
22+
23+
When compaction is triggered:
24+
1. **Preserve system messages** - Always kept intact
25+
2. **Keep recent messages** - Last N messages (default: 10) stay unchanged
26+
3. **Summarize old messages** - Use LLM to create a concise summary
27+
4. **Replace history** - Old messages → summary message + recent messages
28+
29+
### Example Flow
30+
31+
**Before compaction** (120K tokens, 92% of 128K limit):
32+
```
33+
[System] You are a helpful assistant.
34+
[User] Message 1
35+
[Assistant] Response 1
36+
... (100 more messages) ...
37+
[User] Message 102
38+
[Assistant] Response 102
39+
```
40+
41+
**After compaction** (~30K tokens):
42+
```
43+
[System] You are a helpful assistant.
44+
[System] [Conversation Summary] The user asked about... The assistant helped with...
45+
[User] Message 93
46+
[Assistant] Response 93
47+
... (last 10 messages) ...
48+
[User] Message 102
49+
[Assistant] Response 102
50+
```
51+
52+
## Configuration Parameters
53+
54+
| Parameter | Type | Default | Description |
55+
|-----------|------|---------|-------------|
56+
| `auto_compact_enabled` | bool | `True` | Enable/disable auto-compact |
57+
| `auto_compact_threshold` | float | `0.92` | Trigger at X% of context window |
58+
| `auto_compact_keep_recent` | int | `10` | Keep last N messages unchanged |
59+
| `default_context_window` | int | `128000` | Default context size (128K tokens) |
60+
| `compact_model` | str | `None` | Model for summarization, None uses agent's LLM |
61+
62+
## Usage Examples
63+
64+
### Basic Usage
65+
66+
```python
67+
from minion.agents.code_agent import CodeAgent
68+
69+
# Uses default config (auto-compact enabled)
70+
agent = await CodeAgent.create(
71+
name="My Agent",
72+
llm="gpt-4o",
73+
)
74+
```
75+
76+
### Custom Configuration
77+
78+
```python
79+
# More aggressive compaction for smaller context models
80+
agent = await CodeAgent.create(
81+
name="My Agent",
82+
llm="gpt-4o-mini",
83+
auto_compact_enabled=True,
84+
auto_compact_threshold=0.80, # Trigger at 80%
85+
auto_compact_keep_recent=5, # Keep only 5 recent messages
86+
compact_model="gpt-4o-mini", # Use cheaper model for summarization
87+
)
88+
```
89+
90+
### Disable Auto-compact
91+
92+
```python
93+
# Disable (for debugging or when using very large context models)
94+
agent = await CodeAgent.create(
95+
name="My Agent",
96+
llm="gpt-4o",
97+
auto_compact_enabled=False,
98+
)
99+
```
100+
101+
### Manual Compaction
102+
103+
```python
104+
# Trigger compaction manually at any time
105+
await agent.compact_now()
106+
107+
# Or with specific state
108+
await agent.compact_now(state=my_state)
109+
```
110+
111+
## Best Practices
112+
113+
### 1. Threshold Settings
114+
115+
- **Small context models (8K-32K)**: `auto_compact_threshold=0.70-0.80`
116+
- **Medium context models (128K)**: `auto_compact_threshold=0.85-0.92` (default)
117+
- **Large context models (200K+)**: `auto_compact_threshold=0.90-0.95`
118+
119+
### 2. Keep Recent Settings
120+
121+
- **Fast-paced conversations**: `auto_compact_keep_recent=5-8`
122+
- **Standard tasks**: `auto_compact_keep_recent=10` (default)
123+
- **Tasks requiring more context**: `auto_compact_keep_recent=15-20`
124+
125+
### 3. Compact Model Selection
126+
127+
Using a cheaper/faster model for summarization can reduce costs:
128+
129+
```python
130+
agent = await CodeAgent.create(
131+
llm="gpt-4o", # Main reasoning model
132+
compact_model="gpt-4o-mini", # Cheaper model for summaries
133+
)
134+
```
135+
136+
### 4. Combined with Auto-decay
137+
138+
Auto-compact and Auto-decay work together for multi-level context management:
139+
140+
```python
141+
agent = await CodeAgent.create(
142+
# Auto-decay: Large single tool outputs -> files
143+
decay_enabled=True,
144+
decay_ttl_steps=3,
145+
decay_min_size=100_000,
146+
147+
# Auto-compact: Overall history compression
148+
auto_compact_enabled=True,
149+
auto_compact_threshold=0.92,
150+
)
151+
```
152+
153+
**Processing order**:
154+
1. Decay first (save large outputs to files)
155+
2. Then check if compact is needed (compress overall history)
156+
157+
## Technical Details
158+
159+
### Token Calculation
160+
161+
Token count is calculated using the `tiktoken` library with model-specific encoding:
162+
- Uses actual model encoding when available
163+
- Falls back to `cl100k_base` encoding for unknown models
164+
165+
### Context Window Detection
166+
167+
The system automatically detects context window size:
168+
1. Looks up model in known model database
169+
2. Falls back to `default_context_window` if unknown
170+
171+
### Summary Message Format
172+
173+
The summary is inserted as a system message:
174+
```python
175+
{
176+
"role": "system",
177+
"content": "[Conversation Summary]\n...\n\n[End of Summary - Recent messages follow]"
178+
}
179+
```
180+
181+
## FAQ
182+
183+
### Q: Will important information be lost during compaction?
184+
185+
A: The LLM generates a comprehensive summary that captures key information, decisions, and context. Recent messages are always preserved unchanged.
186+
187+
### Q: Can I see when compaction happens?
188+
189+
A: Yes, check the logs:
190+
```
191+
INFO - Auto compact triggered: 118000 tokens >= 117760 (92% of 128000)
192+
INFO - Manual compaction completed
193+
```
194+
195+
### Q: Does compaction affect tool call history?
196+
197+
A: Tool calls are summarized along with other messages. The agent can still access original tool outputs if they were saved by auto-decay.
198+
199+
### Q: What if the summary LLM call fails?
200+
201+
A: The system gracefully falls back to keeping the original history unchanged.
202+
203+
## Related Documentation
204+
205+
- [Auto-decay Guide](auto_decay.md) - Managing large tool responses
206+
- [CodeAgent Documentation](merged_code_agent.md) - Complete CodeAgent documentation

0 commit comments

Comments
 (0)