Overview
Streaming lets you receive partial output as it is generated, which improves perceived latency and user experience. For new OpenAI-style integrations whose selected model exposes a native Responses endpoint, prefer Responses streaming. Responses availability requires the model details to advertise that request format and a same-protocol route to be currently available. If your framework or model uses Chat Completions streaming, AI Sonar supports that compatibility path too.Recommended: Responses Streaming
Responses WebSocket
Connect towss://api.aisonar.dev/v1/responses and send official response.create events. The WebSocket surface does not expose response.cancel; stream is implicit and background is unsupported on this transport. A response.create with generate: false creates a warmup response with no model output or charge and returns an ID that can be continued.
Each connection processes one active response at a time, does not multiplex responses, and has a 60-minute maximum lifetime. Gateway-created response events use a response-scoped monotonic sequence_number. Stream events named error use the flat event shape; connection and protocol errors retain their nested error object.
Chat Completions Streaming
If your framework still expects SSE chunks from/v1/chat/completions, that also works:
Gemini Streaming
POST /v1beta/models/{model}:streamGenerateContent?alt=sse preserves Gemini-native chunks. Metadata-only events, intermediate chunks without finishReason, and natural EOF are valid. AI Sonar does not append a Chat Completions [DONE] marker.
Stream End Conditions
Typical completion conditions:response.completedfor Responses API streamsfinish_reason: "stop"for Chat Completions streamsfinish_reason: "length"when a token limit is hit- tool/function call events when the model wants to use tools
Web App Pattern
Best Practices
Prefer Responses streaming for new builds
Prefer Responses streaming for new builds
Use
/v1/responses if your SDK or app already supports it. Keep /v1/chat/completions streaming for compatibility-driven integrations.Flush output incrementally
Flush output incrementally
Append delta chunks to the UI or terminal as they arrive rather than waiting for the full response.
Handle disconnects and retries
Handle disconnects and retries
Treat network and service disconnects as normal failure modes and reconnect carefully for long-running sessions.