|
| 1 | +# Auto-compact: Automatic Context Window Management |
| 2 | + |
| 3 | +## Overview |
| 4 | + |
| 5 | +As agent conversations grow longer, the context window fills up with historical messages. **Auto-compact** automatically manages context by summarizing older messages when the conversation approaches the context limit. |
| 6 | + |
| 7 | +**Auto-compact** helps prevent: |
| 8 | +- Context overflow errors |
| 9 | +- Performance degradation from processing excessive tokens |
| 10 | +- Increased API costs from large context windows |
| 11 | + |
| 12 | +## How It Works |
| 13 | + |
| 14 | +### 1. Monitor Phase |
| 15 | + |
| 16 | +After each step, the system monitors context usage: |
| 17 | +- Calculates current token count |
| 18 | +- Compares against the model's context window limit |
| 19 | +- Triggers compaction when usage exceeds threshold (default: 92%) |
| 20 | + |
| 21 | +### 2. Compact Phase |
| 22 | + |
| 23 | +When compaction is triggered: |
| 24 | +1. **Preserve system messages** - Always kept intact |
| 25 | +2. **Keep recent messages** - Last N messages (default: 10) stay unchanged |
| 26 | +3. **Summarize old messages** - Use LLM to create a concise summary |
| 27 | +4. **Replace history** - Old messages → summary message + recent messages |
| 28 | + |
| 29 | +### Example Flow |
| 30 | + |
| 31 | +**Before compaction** (120K tokens, 92% of 128K limit): |
| 32 | +``` |
| 33 | +[System] You are a helpful assistant. |
| 34 | +[User] Message 1 |
| 35 | +[Assistant] Response 1 |
| 36 | +... (100 more messages) ... |
| 37 | +[User] Message 102 |
| 38 | +[Assistant] Response 102 |
| 39 | +``` |
| 40 | + |
| 41 | +**After compaction** (~30K tokens): |
| 42 | +``` |
| 43 | +[System] You are a helpful assistant. |
| 44 | +[System] [Conversation Summary] The user asked about... The assistant helped with... |
| 45 | +[User] Message 93 |
| 46 | +[Assistant] Response 93 |
| 47 | +... (last 10 messages) ... |
| 48 | +[User] Message 102 |
| 49 | +[Assistant] Response 102 |
| 50 | +``` |
| 51 | + |
| 52 | +## Configuration Parameters |
| 53 | + |
| 54 | +| Parameter | Type | Default | Description | |
| 55 | +|-----------|------|---------|-------------| |
| 56 | +| `auto_compact_enabled` | bool | `True` | Enable/disable auto-compact | |
| 57 | +| `auto_compact_threshold` | float | `0.92` | Trigger at X% of context window | |
| 58 | +| `auto_compact_keep_recent` | int | `10` | Keep last N messages unchanged | |
| 59 | +| `default_context_window` | int | `128000` | Default context size (128K tokens) | |
| 60 | +| `compact_model` | str | `None` | Model for summarization, None uses agent's LLM | |
| 61 | + |
| 62 | +## Usage Examples |
| 63 | + |
| 64 | +### Basic Usage |
| 65 | + |
| 66 | +```python |
| 67 | +from minion.agents.code_agent import CodeAgent |
| 68 | + |
| 69 | +# Uses default config (auto-compact enabled) |
| 70 | +agent = await CodeAgent.create( |
| 71 | + name="My Agent", |
| 72 | + llm="gpt-4o", |
| 73 | +) |
| 74 | +``` |
| 75 | + |
| 76 | +### Custom Configuration |
| 77 | + |
| 78 | +```python |
| 79 | +# More aggressive compaction for smaller context models |
| 80 | +agent = await CodeAgent.create( |
| 81 | + name="My Agent", |
| 82 | + llm="gpt-4o-mini", |
| 83 | + auto_compact_enabled=True, |
| 84 | + auto_compact_threshold=0.80, # Trigger at 80% |
| 85 | + auto_compact_keep_recent=5, # Keep only 5 recent messages |
| 86 | + compact_model="gpt-4o-mini", # Use cheaper model for summarization |
| 87 | +) |
| 88 | +``` |
| 89 | + |
| 90 | +### Disable Auto-compact |
| 91 | + |
| 92 | +```python |
| 93 | +# Disable (for debugging or when using very large context models) |
| 94 | +agent = await CodeAgent.create( |
| 95 | + name="My Agent", |
| 96 | + llm="gpt-4o", |
| 97 | + auto_compact_enabled=False, |
| 98 | +) |
| 99 | +``` |
| 100 | + |
| 101 | +### Manual Compaction |
| 102 | + |
| 103 | +```python |
| 104 | +# Trigger compaction manually at any time |
| 105 | +await agent.compact_now() |
| 106 | + |
| 107 | +# Or with specific state |
| 108 | +await agent.compact_now(state=my_state) |
| 109 | +``` |
| 110 | + |
| 111 | +## Best Practices |
| 112 | + |
| 113 | +### 1. Threshold Settings |
| 114 | + |
| 115 | +- **Small context models (8K-32K)**: `auto_compact_threshold=0.70-0.80` |
| 116 | +- **Medium context models (128K)**: `auto_compact_threshold=0.85-0.92` (default) |
| 117 | +- **Large context models (200K+)**: `auto_compact_threshold=0.90-0.95` |
| 118 | + |
| 119 | +### 2. Keep Recent Settings |
| 120 | + |
| 121 | +- **Fast-paced conversations**: `auto_compact_keep_recent=5-8` |
| 122 | +- **Standard tasks**: `auto_compact_keep_recent=10` (default) |
| 123 | +- **Tasks requiring more context**: `auto_compact_keep_recent=15-20` |
| 124 | + |
| 125 | +### 3. Compact Model Selection |
| 126 | + |
| 127 | +Using a cheaper/faster model for summarization can reduce costs: |
| 128 | + |
| 129 | +```python |
| 130 | +agent = await CodeAgent.create( |
| 131 | + llm="gpt-4o", # Main reasoning model |
| 132 | + compact_model="gpt-4o-mini", # Cheaper model for summaries |
| 133 | +) |
| 134 | +``` |
| 135 | + |
| 136 | +### 4. Combined with Auto-decay |
| 137 | + |
| 138 | +Auto-compact and Auto-decay work together for multi-level context management: |
| 139 | + |
| 140 | +```python |
| 141 | +agent = await CodeAgent.create( |
| 142 | + # Auto-decay: Large single tool outputs -> files |
| 143 | + decay_enabled=True, |
| 144 | + decay_ttl_steps=3, |
| 145 | + decay_min_size=100_000, |
| 146 | + |
| 147 | + # Auto-compact: Overall history compression |
| 148 | + auto_compact_enabled=True, |
| 149 | + auto_compact_threshold=0.92, |
| 150 | +) |
| 151 | +``` |
| 152 | + |
| 153 | +**Processing order**: |
| 154 | +1. Decay first (save large outputs to files) |
| 155 | +2. Then check if compact is needed (compress overall history) |
| 156 | + |
| 157 | +## Technical Details |
| 158 | + |
| 159 | +### Token Calculation |
| 160 | + |
| 161 | +Token count is calculated using the `tiktoken` library with model-specific encoding: |
| 162 | +- Uses actual model encoding when available |
| 163 | +- Falls back to `cl100k_base` encoding for unknown models |
| 164 | + |
| 165 | +### Context Window Detection |
| 166 | + |
| 167 | +The system automatically detects context window size: |
| 168 | +1. Looks up model in known model database |
| 169 | +2. Falls back to `default_context_window` if unknown |
| 170 | + |
| 171 | +### Summary Message Format |
| 172 | + |
| 173 | +The summary is inserted as a system message: |
| 174 | +```python |
| 175 | +{ |
| 176 | + "role": "system", |
| 177 | + "content": "[Conversation Summary]\n...\n\n[End of Summary - Recent messages follow]" |
| 178 | +} |
| 179 | +``` |
| 180 | + |
| 181 | +## FAQ |
| 182 | + |
| 183 | +### Q: Will important information be lost during compaction? |
| 184 | + |
| 185 | +A: The LLM generates a comprehensive summary that captures key information, decisions, and context. Recent messages are always preserved unchanged. |
| 186 | + |
| 187 | +### Q: Can I see when compaction happens? |
| 188 | + |
| 189 | +A: Yes, check the logs: |
| 190 | +``` |
| 191 | +INFO - Auto compact triggered: 118000 tokens >= 117760 (92% of 128000) |
| 192 | +INFO - Manual compaction completed |
| 193 | +``` |
| 194 | + |
| 195 | +### Q: Does compaction affect tool call history? |
| 196 | + |
| 197 | +A: Tool calls are summarized along with other messages. The agent can still access original tool outputs if they were saved by auto-decay. |
| 198 | + |
| 199 | +### Q: What if the summary LLM call fails? |
| 200 | + |
| 201 | +A: The system gracefully falls back to keeping the original history unchanged. |
| 202 | + |
| 203 | +## Related Documentation |
| 204 | + |
| 205 | +- [Auto-decay Guide](auto_decay.md) - Managing large tool responses |
| 206 | +- [CodeAgent Documentation](merged_code_agent.md) - Complete CodeAgent documentation |
0 commit comments