Skip to content

test: add reusable end-to-end LLM mock server - #133

Closed
nachiketb-nvidia wants to merge 2 commits into
mainfrom
nachiketb/e2e-mock-server
Closed

nachiketb-nvidia wants to merge 2 commits into
mainfrom
nachiketb/e2e-mock-server

Conversation

@nachiketb-nvidia

@nachiketb-nvidia nachiketb-nvidia commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

What

  • add a private, non-publishable switchyard-test-server workspace crate
  • serve OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages
  • support buffered and SSE streaming responses through the existing translation codecs
  • provide a curated random response bank plus deterministic per-test overrides
  • support model/header-based HTTP error injection and request capture
  • expose both an in-process fixture and a standalone binary
  • replace the inline mock implementation in switchyard-server integration tests

Why

Switchyard integration tests need one reusable provider mock rather than independent,
partially compatible servers embedded in each test suite. Generating responses through
switchyard-translation keeps the mock aligned with the same wire formats the product
supports.

Review

  • confirm the fixture API is intentionally small and sufficient for server/client tests
  • confirm the built-in response bank covers useful text shapes without making exact tests flaky
  • confirm publish = false and dev-only consumption keep this out of production artifacts

Validation

  • cargo test -p switchyard-test-server
  • cargo test -p switchyard-server
  • cargo test --workspace
  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • standalone binary HTTP smoke against /v1/chat/completions

Summary by CodeRabbit

  • New Features

    • Added a reusable mock LLM server for integration testing across OpenAI and Anthropic-compatible endpoints.
    • Supports configurable responses, streaming output, request capture, and simulated errors.
    • Added a standalone mode for running the mock server locally.
  • Tests

    • Updated server integration tests to use the shared mock server.
    • Added coverage for buffered and streaming responses, request recording, default responses, and error handling.
  • Documentation

    • Added guidance recommending the shared mock server for Rust HTTP end-to-end tests.

Signed-off-by: nachiketb <nachiketb@nvidia.com>
@nachiketb-nvidia
nachiketb-nvidia marked this pull request as ready for review July 23, 2026 23:08
@nachiketb-nvidia
nachiketb-nvidia requested a review from a team as a code owner July 23, 2026 23:08
Signed-off-by: nachiketb <nachiketb@nvidia.com>
@coderabbitai

coderabbitai Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Changes

Mock LLM Server

Layer / File(s) Summary
Workspace integration
Cargo.toml, crates/switchyard-test-server/Cargo.toml, crates/switchyard-server/Cargo.toml, .agents/skills/...
Registers the private crate, connects it to server tests, and documents its use for Rust HTTP end-to-end testing.
Mock server runtime
crates/switchyard-test-server/src/lib.rs, crates/switchyard-test-server/src/main.rs
Adds configurable provider endpoints, request capture, buffered and streaming responses, error injection, and a standalone executable.
Integration test adoption
crates/switchyard-server/tests/server.rs, crates/switchyard-test-server/tests/server.rs
Migrates existing tests to the reusable server and adds coverage for responses, streaming, defaults, captures, and errors.

Estimated code review effort: 4 (Complex) | ~45 minutes

Poem

I’m a bunny with a server to run,
Mocking each model in the sun.
Streams hop by, errors too,
Captured requests say “peekaboo!”
Tests now dance in a carrot queue.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 48.39% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately summarizes the main change: adding a reusable end-to-end LLM mock server.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
crates/switchyard-test-server/tests/server.rs (1)

16-155: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add concise comments for the protocol behaviors under test.

The test names identify broad intent, but the provider-specific response, SSE framing, and canonical-error contracts are important enough to document inline.

  • crates/switchyard-test-server/tests/server.rs#L16-L155: add short comments describing the buffered-format, stream-terminator, and error-injection contracts.
  • crates/switchyard-server/tests/server.rs#L116-L188: document why classifier routing produces three upstream calls.
  • crates/switchyard-server/tests/server.rs#L316-L387: document the expected OpenAI stream terminator and canonical upstream-error mapping.

As per coding guidelines, use concise comments for important behavioral tests.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/switchyard-test-server/tests/server.rs` around lines 16 - 155, Add
concise inline comments at crates/switchyard-test-server/tests/server.rs lines
16-155 documenting buffered provider-format responses, SSE stream
terminators/framing, and model/header error injection. At
crates/switchyard-server/tests/server.rs lines 116-188, explain why classifier
routing results in three upstream calls; at lines 316-387, document the expected
OpenAI stream terminator and canonical upstream-error mapping. Modify only these
behavioral-test comments.

Source: Coding guidelines

crates/switchyard-test-server/src/lib.rs (1)

224-264: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Document the request and streaming behavior contracts.

Add brief comments for capture-before-error handling, model-error precedence, and codec/SSE terminal framing. These are non-obvious test-server contracts.

Suggested clarification
 async fn handle_request(...) -> Response {
+    // Capture before applying injected failures so tests can inspect every request.
     state.requests.lock().await.push(...);

+    // Model-specific failures take precedence over the per-request status override.
     if let Some(status) = state.model_errors.get(model).copied() {
 fn stream_response(...) -> Response {
+    // Translate canonical chunks, then append OpenAI's required `[DONE]` terminator.
     let chunks: LlmResponseStream = Box::pin(...);

As per coding guidelines, use concise comments for complex validation, routing, async, and lifecycle logic.

Also applies to: 337-366

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/switchyard-test-server/src/lib.rs` around lines 224 - 264, Add concise
comments in handle_request documenting that requests are captured before any
error response, configured model errors take precedence over header-requested
errors, and streaming responses must preserve codec-specific SSE terminal
framing. Add the corresponding terminal-framing comment in stream_response,
without changing behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@crates/switchyard-test-server/src/lib.rs`:
- Around line 224-264: Add concise comments in handle_request documenting that
requests are captured before any error response, configured model errors take
precedence over header-requested errors, and streaming responses must preserve
codec-specific SSE terminal framing. Add the corresponding terminal-framing
comment in stream_response, without changing behavior.

In `@crates/switchyard-test-server/tests/server.rs`:
- Around line 16-155: Add concise inline comments at
crates/switchyard-test-server/tests/server.rs lines 16-155 documenting buffered
provider-format responses, SSE stream terminators/framing, and model/header
error injection. At crates/switchyard-server/tests/server.rs lines 116-188,
explain why classifier routing results in three upstream calls; at lines
316-387, document the expected OpenAI stream terminator and canonical
upstream-error mapping. Modify only these behavioral-test comments.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: acfad70e-41db-4cbe-949d-5ad6788bfa4a

📥 Commits

Reviewing files that changed from the base of the PR and between 1bafb9b and b4e5daf.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock, !Cargo.lock
📒 Files selected for processing (8)
  • .agents/skills/switchyard-testing-ci/SKILL.md
  • Cargo.toml
  • crates/switchyard-server/Cargo.toml
  • crates/switchyard-server/tests/server.rs
  • crates/switchyard-test-server/Cargo.toml
  • crates/switchyard-test-server/src/lib.rs
  • crates/switchyard-test-server/src/main.rs
  • crates/switchyard-test-server/tests/server.rs

@grahamking

grahamking commented Jul 24, 2026 •

Copy link
Copy Markdown
Contributor

Could we try https://github.com/vidaiUK/VidaiMock or https://github.com/StacklokLabs/mockllm instead, to avoid maintaining it ourselves, and get much higher API coverage?

In my experience these tend to turn into a full time job.

@nachiketb-nvidia
nachiketb-nvidia deleted the nachiketb/e2e-mock-server branch July 24, 2026 18:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants