Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,19 @@ All notable changes to NadirClaw will be documented in this file.

## [Unreleased]

## [0.20.0] - 2026-06-12

### Added
- **Context-optimizer compression upgrades** (#65):
- **Pluggable backend** — `NADIRCLAW_OPTIMIZE_BACKEND` selects `native` (default, built-in stdlib pipeline) or `headroom` (opt-in, delegates to the Apache-2.0 [`headroom-ai`](https://pypi.org/project/headroom-ai/) package via `pip install nadirclaw[headroom]`). Headroom is lazy and **fail-open**: if it is not installed or raises, the optimizer transparently falls back to `native` and the request never fails. Per-request override via `optimize_backend` in the body.
- **Progressive (staged) compression** — `--optimize progressive` / `NADIRCLAW_OPTIMIZE=progressive` runs an escalation ladder (`native_safe → native_aggressive → headroom_structural → headroom_ml`) that **stops as soon as `NADIRCLAW_OPTIMIZE_TARGET_TOKENS` is met**. With no budget set it stops after `native_aggressive` (dependency-free, lossless); Headroom stages are skipped silently when `headroom-ai` is absent; the lossy ML stage runs only when `NADIRCLAW_OPTIMIZE_ALLOW_LOSSY` is on. Tunable via `NADIRCLAW_OPTIMIZE_MAX_STAGE`. New library entrypoint `nadirclaw.optimize.compress_progressive()`.
- **Columnar JSON-array packing** (`json_array_pack`, aggressive mode) — rewrites homogeneous arrays of same-keyed objects (DB results, API list responses, large tool outputs) into a header plus one value-array per row, emitting each key once. Information-lossless and deterministically reversible; ~68% vs pretty-printed JSON. Never runs in `safe` mode.
- **Native CCR** (`nadirclaw/ccr.py`) — deterministic offload + `nadir_retrieve` fetch-back loop that moves oversized content out of the prompt behind a retrieve handle, fully reversible because the originals are kept server-side. Library-only for now (not yet wired into `nadirclaw serve`).
- Apache-2.0 attribution for `headroom-ai` in `THIRD_PARTY_NOTICES.md`; docs, benchmarks, and tests (ccr, progressive, json_array_pack, backends, code-safety).

### Fixed
- **Whitespace normalization corrupted unfenced source code** — the `whitespace_normalize` transform collapsed the *leading indentation* of raw (unfenced) code arriving as file-read tool outputs, flattening nested Python/YAML/diffs into invalid syntax while reporting it as "savings". It now preserves leading indentation and only collapses interior multi-spaces, in both `safe` and `aggressive` modes (#65).

## [0.19.3] - 2026-06-12

### Fixed
Expand Down
48 changes: 46 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,7 +156,7 @@ NadirClaw is the free, open-source core. If you are routing production traffic o

## Features

- **Context Optimize** — compacts bloated context (JSON, tool schemas, chat history, whitespace) before dispatch, saving 30-70% input tokens with zero semantic loss. Modes: `off` (default), `safe` (lossless), `aggressive` (future). See [savings analysis](docs/context-optimize-savings.md)
- **Context Optimize** — compacts bloated context (JSON, tool schemas, chat history, whitespace) before dispatch, saving 30-70% input tokens with zero semantic loss. Modes: `off` (default), `safe` (lossless), `aggressive` (+ columnar JSON-array packing & semantic dedup), `progressive` (staged ladder that only escalates until a token budget is met). Pluggable backend (`native` default, or opt-in `headroom`). See [savings analysis](docs/context-optimize-savings.md)
- **Smart routing** — classifies prompts in ~10ms using sentence embeddings
- **Pluggable classifier** — `binary` (default, ~10ms centroid classifier) or `distilbert` (3-class fine-tuned DistilBERT that natively predicts simple/mid/complex). Select with `NADIRCLAW_COMPLEXITY_ANALYZER`
- **Three-tier routing** — simple / mid / complex tiers with configurable score thresholds (`NADIRCLAW_TIER_THRESHOLDS`); set `NADIRCLAW_MID_MODEL` for a cost-effective middle tier
Expand Down Expand Up @@ -1047,7 +1047,7 @@ Test context compaction on a file or stdin without running the server:
```bash
nadirclaw optimize payload.json # dry-run with safe mode
nadirclaw optimize payload.json --format json # machine-readable output
nadirclaw optimize payload.json --mode aggressive # aggressive mode (future)
nadirclaw optimize payload.json --mode aggressive # + columnar JSON packing & semantic dedup
cat messages.json | nadirclaw optimize # pipe from stdin
```

Expand Down Expand Up @@ -1375,6 +1375,50 @@ On top of routing savings, Context Optimize compacts bloated payloads before the

Average: **61.5% input token reduction** across structured payloads. Enable with `--optimize safe`. See [full analysis](docs/context-optimize-savings.md).

#### Modes

```bash
# Pick a mode when starting the server (or set NADIRCLAW_OPTIMIZE)
nadirclaw serve --optimize safe # lossless: dedup, json minify, whitespace
nadirclaw serve --optimize aggressive # + columnar JSON packing & semantic dedup
nadirclaw serve --optimize progressive # staged ladder, escalates only until budget met
```

Per request, override the mode in the body: `{"optimize": "aggressive", "messages": [...]}` (or `"off"` to disable).

#### Backends — `native` (default) vs `headroom`

The mode decides *how hard* to compress; the backend decides *who* runs it. `native` is the
built-in, dependency-free pipeline. `headroom` delegates to the optional Apache-2.0
[`headroom-ai`](https://pypi.org/project/headroom-ai/) package for statistical JSON crushing and
content-type routing — and **transparently falls back to `native`** if it isn't installed or errors,
so it never breaks a request.

```bash
pip install "nadirclaw[headroom]"
NADIRCLAW_OPTIMIZE=safe NADIRCLAW_OPTIMIZE_BACKEND=headroom nadirclaw serve
# per-request: {"optimize": "safe", "optimize_backend": "headroom", "messages": [...]}
```

#### Progressive (staged) compression

`progressive` escalates through stages — `native_safe → native_aggressive → headroom_structural →
headroom_ml` — and **stops as soon as a token budget is met**, so you only pay the cost (and
fidelity risk) of heavier compression when lighter stages aren't enough. With no
`NADIRCLAW_OPTIMIZE_TARGET_TOKENS` set, it stops after `native_aggressive` (dependency-free,
lossless); Headroom stages are skipped silently if `headroom-ai` is absent, and the lossy ML stage
only runs when explicitly allowed.

```bash
NADIRCLAW_OPTIMIZE=progressive \
NADIRCLAW_OPTIMIZE_TARGET_TOKENS=180000 \
NADIRCLAW_OPTIMIZE_MAX_STAGE=headroom_structural \
nadirclaw serve
```

See the [backends & progressive reference](docs/context-optimize-savings.md#backends-native-default-vs-headroom)
for the full ladder, the env-var table under [Configuration Reference](#configuration-reference), and safety/fallback details.

## API Endpoints

Auth is disabled by default (local-only). Set `NADIRCLAW_AUTH_TOKEN` to require a bearer token.
Expand Down
2 changes: 1 addition & 1 deletion nadirclaw/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""NadirClaw — Open-source LLM router."""

__version__ = "0.19.3"
__version__ = "0.20.0"
Loading