Skip to content

Commit f9e2fc5

Browse files
maralbahariclaude
andauthored
feat: add Codex-compatible GET v1/models and WebSocket session bug fixes (vllm-project#79)
## Summary `handler.rs` had grown to 630 lines mixing WebSocket logic, HTTP handlers, proxy utilities, and model transformation code in a single file. This PR splits it into a focused module tree and adds a Codex CLI-compatible `/v1/models` endpoint. **`GET /v1/models`: Codex CLI compatibility** - Added `proxy_get(path, headers, state)` to `agentic-core/proxy.rs`: a GET variant of `proxy_request` that applies the same header filtering and auth injection. - Codex CLI is detected via `?client_version=<ver>` query param (typed `ModelsParams` extractor). - Non-Codex clients: upstream vLLM response streamed back unchanged via `proxy_get`. - Codex clients: vLLM `{ "object": "list", "data": [...] }` transformed to `{ "models": [...] }` with the full `ModelInfo` shape Codex expects (`slug`, `context_window`, `auto_review_model_override`, `apply_patch_tool_type`, capabilities, etc.). - Static `ModelInfo` fields built once via `OnceLock` and cloned per model; only the five per-model fields are patched at response time. **WebSocket session bug fixes** Identified and fixed two bugs in `websocket/responses.rs` that prevented Codex CLI from persisting history to the database: - **Pipelined requests caused connection resets**: Codex sends the next `response.create` on the same WebSocket connection while the current stream is still active. The old handler returned a `ConcurrentMessage` error and closed the connection, forcing Codex to reconnect on every turn. Fixed by introducing a `VecDeque` queue: incoming requests that arrive mid-stream are enqueued and processed in order after the current stream completes. The `ConcurrentMessage` error variant was removed as it is no longer reachable. - **`store: false` bypassed the database**: Codex CLI explicitly sends `"store": false` in every request body, which caused the gateway to skip persistence entirely and proxy straight to vLLM. Fixed by forcing `payload.store = true` on the WebSocket path; the gateway is the stateful layer and should always persist regardless of what the client sends. Together these fixes ensure every completed Codex turn is written to the DB with its full SSE history, and that multi-turn conversations chain correctly via `previous_response_id`. **Handler reorganization** - `handler.rs` deleted. Replaced by `handler/` module tree: - `common.rs`: shared utilities used across handlers: `convert_response`, `executor_error_response`, `read_bytes`, `resolve_exec_ctx`, `sse_response` - `http/conversations.rs`: `POST /v1/conversations` - `http/models.rs`: `GET /health`, `GET /ready`, `GET /v1/models` + Codex model transform logic - `http/responses.rs`: `POST /v1/responses` (proxy and stateful paths) - `websocket/responses.rs`: WebSocket `/v1/responses` handler and streaming loop - `websocket/error.rs`: `WsError` enum with status/code/frame helpers - `handler/mod.rs` re-exports all public handlers; no import paths in `app.rs` changed. ## Codex CLI setup Create `$CODEX_HOME/config.toml` (e.g. `~/.codex/config.toml`) pointing at agentic-api: ```toml model_provider = "agentic-api" [model_providers.agentic-api] name = "agentic-api" base_url = "http://localhost:9000/v1" wire_api = "responses" requires_openai_auth = false supports_websockets = true ``` Run with: ```bash codex --disable image_generation -c model_provider=agentic-api -m "model_name" ``` ## Test Plan - `cargo clippy --all-targets -- -D warnings` clean - `cargo test`: all tests pass, 0 failed --------- Signed-off-by: maral <maralbahari.98@gmail.com> Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
1 parent 647e49d commit f9e2fc5

12 files changed

Lines changed: 761 additions & 500 deletions

File tree

‎crates/agentic-core/src/proxy.rs‎

Lines changed: 42 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -153,6 +153,48 @@ pub fn error_response(status: StatusCode, code: &str, message: &str) -> ProxyRes
153153
}
154154
}
155155

156+
/// Proxy a GET request to an arbitrary upstream path.
157+
///
158+
/// Applies the same header filtering and auth injection as [`proxy_request`].
159+
/// Uses the non-streaming client; the response body is returned as a full
160+
/// [`ProxyBody::Full`] payload.
161+
pub async fn proxy_get(path: &str, request_headers: &HeaderMap, state: &ProxyState) -> ProxyResponse {
162+
let llm_headers = filter_request_headers(request_headers, &state.config);
163+
let base = state.config.llm_api_base.trim_end_matches('/');
164+
let url = format!("{base}/{}", path.trim_start_matches('/'));
165+
166+
let llm_resp = match state.non_stream_client.get(&url).headers(llm_headers).send().await {
167+
Ok(r) => r,
168+
Err(e) if e.is_timeout() => {
169+
warn!("upstream GET {path} timed out: {e}");
170+
return error_response(StatusCode::GATEWAY_TIMEOUT, "upstream_timeout", "upstream timeout");
171+
}
172+
Err(e) => {
173+
warn!("upstream GET {path} failed: {e}");
174+
return error_response(StatusCode::BAD_GATEWAY, "upstream_unavailable", "upstream unavailable");
175+
}
176+
};
177+
178+
let status = StatusCode::from_u16(llm_resp.status().as_u16()).unwrap_or(StatusCode::BAD_GATEWAY);
179+
let response_headers = filter_response_headers(llm_resp.headers());
180+
181+
match llm_resp.bytes().await {
182+
Ok(payload) => ProxyResponse {
183+
status,
184+
headers: response_headers,
185+
body: ProxyBody::Full(payload),
186+
},
187+
Err(e) => {
188+
warn!("failed to read upstream GET {path} body: {e}");
189+
error_response(
190+
StatusCode::BAD_GATEWAY,
191+
"upstream_unavailable",
192+
"failed to read upstream response",
193+
)
194+
}
195+
}
196+
}
197+
156198
pub async fn proxy_request(request: ProxyRequest, state: &ProxyState) -> ProxyResponse {
157199
let is_streaming = serde_json::from_slice::<Value>(&request.body)
158200
.ok()

‎crates/agentic-server/src/app.rs‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -9,7 +9,7 @@ use tower_http::cors::{AllowOrigin, Any, CorsLayer};
99
use agentic_core::executor::ExecutionContext;
1010
use agentic_core::proxy::ProxyState;
1111

12-
use crate::handler::{conversations, health, ready, responses, responses_ws};
12+
use crate::handler::{conversations, health, models, ready, responses, responses_ws};
1313

1414
/// Server-level configuration read from environment variables.
1515
pub struct ServerConfig {
@@ -71,6 +71,7 @@ pub fn build_router(state: AppState, server_config: &ServerConfig) -> Router {
7171
.route("/health", get(health))
7272
.route("/ready", get(ready))
7373
.route("/v1/conversations", post(conversations))
74+
.route("/v1/models", get(models))
7475
.route("/v1/responses", post(responses).get(responses_ws))
7576
.layer(server_config.cors_layer())
7677
.with_state(state)

0 commit comments

Comments
 (0)