Skip to content

Commit b32c4ef

Browse files
jayzuccarelliclaude
andcommitted
broker: REASONING_EFFORT knob for reasoning-line Realtime models
gpt-realtime-2.1 is a reasoning-line model; its server default effort (low) pads spoken replies and adds latency. Pipecat 0.0.97's SessionProperties has no reasoning field and pydantic serializes by declared type, so the broker injects reasoning.effort into the session.update wire dict when REASONING_EFFORT is set. Unset = field not sent (required for non-reasoning models like gpt-realtime); zero behavior change by default. Benched 2026-07-21 (isolated 8766/8767, hygiene off, paired same-hour runs): minimal effort fixes 2.1's verbosity (one-liner replies) but first-audio p50 stays 500-680ms vs 97-292ms on live gpt-realtime, so the live model keeps MODEL=gpt-realtime for now. Part of JAY-92. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1 parent b9ee7f2 commit b32c4ef

3 files changed

Lines changed: 26 additions & 0 deletions

File tree

‎README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -68,6 +68,7 @@ The broker fetches HA's tools at startup and registers them on the Realtime sess
6868
|---|---|---|
6969
| `OPENAI_API_KEY` | — | required |
7070
| `MODEL` | `gpt-realtime` | Realtime model |
71+
| `REASONING_EFFORT` | unset | reasoning effort for reasoning-line models (`gpt-realtime-2.1`+): `minimal`/`low`/`medium`/`high`/`xhigh`; leave unset for non-reasoning models |
7172
| `VOICE` | `marin` | Realtime voice |
7273
| `INSTRUCTIONS` | generic | system prompt / persona |
7374
| `WS_HOST` / `WS_PORT` | `0.0.0.0` / `8765` | where the device connects |

‎broker/realtime_broker/agent.py‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,8 +16,10 @@
1616
AudioConfiguration,
1717
AudioInput,
1818
AudioOutput,
19+
ClientEvent,
1920
InputAudioTranscription,
2021
SessionProperties,
22+
SessionUpdateEvent,
2123
TurnDetection,
2224
)
2325
from pipecat.services.openai.realtime.llm import OpenAIRealtimeLLMService
@@ -42,11 +44,27 @@ class VoicePERealtimeService(OpenAIRealtimeLLMService):
4244
its own response.create).
4345
"""
4446

47+
def __init__(self, *args, reasoning_effort: str | None = None, **kwargs):
48+
super().__init__(*args, **kwargs)
49+
self._reasoning_effort = reasoning_effort
50+
4551
async def _handle_context(self, context: LLMContext) -> None:
4652
self._context = context
4753
self._llm_needs_conversation_setup = False
4854
await self._process_completed_function_calls(send_new_results=True)
4955

56+
async def send_client_event(self, event: ClientEvent) -> None:
57+
# Pipecat 0.0.97's SessionProperties has no `reasoning` field and
58+
# pydantic serializes by declared type, so a subclass field would be
59+
# dropped — inject into the wire dict instead. Applies to every
60+
# session.update (initial setup and mid-session).
61+
if self._reasoning_effort and isinstance(event, SessionUpdateEvent):
62+
dump = event.model_dump(exclude_none=True)
63+
dump["session"]["reasoning"] = {"effort": self._reasoning_effort}
64+
await self._ws_send(dump)
65+
return
66+
await super().send_client_event(event)
67+
5068
# Custom broker tools, registered with handlers by the server.
5169
CUSTOM_TOOLS = [
5270
{
@@ -206,6 +224,7 @@ async def build_agent(config: Config, mcp: MCPClient | None) -> OpenAIRealtimeLL
206224
model=config.model,
207225
session_properties=session,
208226
start_audio_paused=False,
227+
reasoning_effort=config.reasoning_effort,
209228
)
210229

211230
if mcp is not None and tools_schema is not None:

‎broker/realtime_broker/config.py‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,11 @@ class Config:
1414
model: str = "gpt-realtime"
1515
voice: str = "marin"
1616
instructions: str = "You are a helpful voice assistant."
17+
# Reasoning effort for reasoning-line Realtime models (gpt-realtime-2.1+):
18+
# minimal/low/medium/high/xhigh. The server default (low) makes the model
19+
# deliberate and pad its spoken replies; "minimal" suits command-and-control.
20+
# None = field not sent, required for non-reasoning models (gpt-realtime).
21+
reasoning_effort: str | None = None
1722

1823
ws_host: str = "0.0.0.0"
1924
ws_port: int = 8765
@@ -80,6 +85,7 @@ def from_env(cls) -> "Config":
8085
openai_api_key=api_key,
8186
model=os.environ.get("MODEL", "gpt-realtime"),
8287
voice=os.environ.get("VOICE", "marin"),
88+
reasoning_effort=os.environ.get("REASONING_EFFORT") or None,
8389
instructions=os.environ.get("INSTRUCTIONS", cls.instructions),
8490
ws_host=os.environ.get("WS_HOST", "0.0.0.0"),
8591
ws_port=int(os.environ.get("WS_PORT", "8765")),

0 commit comments

Comments
 (0)