> ## Documentation Index
> Fetch the complete documentation index at: https://docs.aisonar.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Implement real-time streaming responses

## Overview

Streaming lets you receive partial output as it is generated, which improves perceived latency and user experience.

For new OpenAI-style integrations whose selected model exposes a native Responses endpoint, prefer **Responses streaming**. Responses availability requires the model details to advertise that request format and a same-protocol route to be currently available. If your framework or model uses Chat Completions streaming, AI Sonar supports that compatibility path too.

## Recommended: Responses Streaming

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.aisonar.dev/v1/responses \
    -H "Authorization: Bearer sk-your-api-key" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "gpt-5.4",
      "input": "Write a short poem.",
      "stream": true
    }'
  ```

  ```python Python theme={null}
  from openai import OpenAI

  client = OpenAI(
      api_key="sk-your-api-key",
      base_url="https://api.aisonar.dev/v1"
  )

  stream = client.responses.create(
      model="gpt-5.4",
      input="Write a short poem.",
      stream=True
  )

  for event in stream:
      if event.type == "response.output_text.delta":
          print(event.delta, end="", flush=True)
  ```

  ```javascript JavaScript theme={null}
  import OpenAI from 'openai';

  const client = new OpenAI({
    apiKey: 'sk-your-api-key',
    baseURL: 'https://api.aisonar.dev/v1'
  });

  const stream = await client.responses.create({
    model: 'gpt-5.4',
    input: 'Write a short poem.',
    stream: true
  });

  for await (const event of stream) {
    if (event.type === 'response.output_text.delta') {
      process.stdout.write(event.delta);
    }
  }
  ```
</CodeGroup>

AI Sonar preserves native Responses SSE event names, order, and fields. It parses usage and terminal state out of band but does not rewrite the wire events. Once the first event has been delivered, the gateway does not retry the request.

## Responses WebSocket

Connect to `wss://api.aisonar.dev/v1/responses` and send official `response.create` events. The WebSocket surface does not expose `response.cancel`; `stream` is implicit and `background` is unsupported on this transport. A `response.create` with `generate: false` creates a warmup response with no model output or charge and returns an ID that can be continued.

Each connection processes one active response at a time, does not multiplex responses, and has a 60-minute maximum lifetime. Gateway-created response events use a response-scoped monotonic `sequence_number`. Stream events named `error` use the flat event shape; connection and protocol errors retain their nested `error` object.

## Chat Completions Streaming

If your framework still expects SSE chunks from `/v1/chat/completions`, that also works:

```python theme={null}
stream = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Write a short poem"}],
    stream=True
)

for chunk in stream:
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)
```

## Gemini Streaming

`POST /v1beta/models/{model}:streamGenerateContent?alt=sse` preserves Gemini-native chunks. Metadata-only events, intermediate chunks without `finishReason`, and natural EOF are valid. AI Sonar does not append a Chat Completions `[DONE]` marker.

## Stream End Conditions

Typical completion conditions:

* `response.completed` for Responses API streams
* `finish_reason: "stop"` for Chat Completions streams
* `finish_reason: "length"` when a token limit is hit
* tool/function call events when the model wants to use tools

## Web App Pattern

```javascript theme={null}
async function streamChat(message) {
  const response = await fetch('https://api.aisonar.dev/v1/chat/completions', {
    method: 'POST',
    headers: {
      'Authorization': 'Bearer sk-your-api-key',
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      model: 'gpt-4o',
      messages: [{ role: 'user', content: message }],
      stream: true
    })
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const chunk = decoder.decode(value);
    const lines = chunk.split('\\n').filter(line => line.startsWith('data: '));

    for (const line of lines) {
      const data = line.slice(6);
      if (data === '[DONE]') return;
      const parsed = JSON.parse(data);
      const content = parsed.choices?.[0]?.delta?.content;
      if (content) {
        document.getElementById('output').textContent += content;
      }
    }
  }
}
```

## Best Practices

<AccordionGroup>
  <Accordion title="Prefer Responses streaming for new builds">
    Use `/v1/responses` if your SDK or app already supports it. Keep `/v1/chat/completions` streaming for compatibility-driven integrations.
  </Accordion>

  <Accordion title="Flush output incrementally">
    Append delta chunks to the UI or terminal as they arrive rather than waiting for the full response.
  </Accordion>

  <Accordion title="Handle disconnects and retries">
    Treat network and service disconnects as normal failure modes and reconnect carefully for long-running sessions.
  </Accordion>
</AccordionGroup>
