# Media streams: bring your own bot

> Stream a call's audio to your own WebSocket server: fork a copy for analytics or recording, or connect your own voice bot in place of MoreVoice's AI while MoreVoice keeps the phone lines, queues, recording and compliance. The protocol is compatible with Twilio Media Streams.

Media streams are part of the Growth, Business, Enterprise and Developer plans.

A **media stream** sends a call's audio, as it happens, to a WebSocket server of yours. It comes in two modes:

- **`fork`**: your server gets a copy of the call's audio (the caller, what the caller hears, or both) and listens: real-time analytics, your own transcription, a compliance recorder. The call goes on as before.
- **`connect`**: your server gets the caller's audio and talks: your bot answers instead of MoreVoice's AI. You send audio back, and can transfer the call, hang up or tag it.

With `connect`, MoreVoice still does everything around the conversation: the phone numbers and SIP trunks, IVR menus and routing, queues and human agents, recording, transcripts and webhooks, and the [§30A compliance](https://docs.morevoice.ai/guides/compliance/) of every outbound call. Your bot does only the talking.

## The protocol

The protocol is **compatible with Twilio Media Streams**: the same JSON messages over one WebSocket, the same field names and the same μ-law audio. A bot written for Twilio works with a URL change, and the bot frameworks that read Twilio streams read these. MoreVoice adds three messages of its own (`transfer`, `hangup` and `metadata`) and linear PCM at 16 kHz, which most speech engines prefer.

Every frame is a JSON text message with an `event` field. In `connect` mode your bot may send messages back; in `fork` mode MoreVoice ignores everything from your server except closing the connection.

| `event` | Direction | What it means | Fields |
| --- | --- | --- | --- |
| `connected` | To your bot | The first frame, once: the protocol's name and version. | `protocol`, `version` |
| `start` | To your bot | The stream's IDs (`accountSid` is your organisation, `callSid` the call, `streamSid` this stream), its tracks, the `customParameters` you set and the audio format. | `sequenceNumber`, `start.accountSid`, `start.streamSid`, `start.callSid`, `start.tracks`, `start.customParameters`, `start.mediaFormat.encoding`, `start.mediaFormat.sampleRate`, `start.mediaFormat.channels`, `start.mediaFormat.byteOrder`, `streamSid` |
| `media` | To your bot | 20 ms of one track's audio, base64, with its chunk number and its time in the stream (ms). | `sequenceNumber`, `media.track`, `media.chunk`, `media.timestamp`, `media.payload`, `streamSid` |
| `dtmf` | To your bot | The caller pressed a key. | `sequenceNumber`, `dtmf.track`, `dtmf.digit`, `streamSid` |
| `mark` | To your bot | Everything you sent before one of your marks has finished playing. | `sequenceNumber`, `mark.name`, `streamSid` |
| `stop` | To your bot | The last frame, once: the stream is over. | `sequenceNumber`, `stop.accountSid`, `stop.callSid`, `streamSid` |
| `media` | From your bot | Audio to play to the caller, base64, in the stream's format. | `streamSid`, `media.payload` |
| `mark` | From your bot | A named point in the audio you sent. It comes back as a `mark` once everything before it has played. | `streamSid`, `mark.name` |
| `clear` | From your bot | Drop the audio still queued to play (the caller interrupted). Pending marks come back at once. | `streamSid` |
| `transfer` | From your bot | Hand the call to a queue (`queue_id`), a phone number (`number`, `cold` or `warm`) or an assistant (`assistant_id`): exactly one. | `streamSid`, `transfer.queue_id`, `transfer.number`, `transfer.assistant_id`, `transfer.mode` |
| `hangup` | From your bot | End the call, with an optional `reason`. | `streamSid`, `hangup.reason` |
| `metadata` | From your bot | Merge key–value strings into the call's `metadata`. | `streamSid`, `metadata` |

### Audio formats

| Format | `mediaFormat.encoding` | Sample rate | Byte order |
| --- | --- | --- | --- |
| `mulaw_8000` (default) | `audio/x-mulaw` | 8,000 Hz | — |
| `l16_16000` | `audio/l16` | 16,000 Hz | little-endian |
| `l16_8000` | `audio/l16` | 8,000 Hz | little-endian |

### Limits

- The `connected` frame names the protocol `Call`, version `1.0.0`.
- A bot may ask for the WebSocket subprotocol `mv-media.v1`; plain connections are accepted too.
- One `media` frame from your bot carries at most 98,304 base64 characters; send long audio as several frames.
- Mark names are echoed back as they are, up to 256 characters.

### Differences from Twilio

Port a Twilio bot by changing the URL, then check these differences:

- **IDs** are MoreVoice's: `accountSid` is your organisation (`org_…`), `callSid` the call (`call_…`, as the API and webhooks show it) and `streamSid` the stream. Treat them as opaque strings.
- **Audio formats.** Besides Twilio's μ-law at 8 kHz, a stream can carry linear 16-bit PCM. MoreVoice's PCM is **little-endian**, what speech engines and bot frameworks read natively, and `mediaFormat.byteOrder` says so (RFC 3551 network order is big-endian): convert if your code assumes otherwise.
- **Tracks** are `inbound` (the caller) and `outbound` (what the caller hears: the assistant, an agent, prompts).
- **More messages.** `transfer`, `hangup` and `metadata` let a `connect` bot finish the job without a separate API call. Ignore any message your bot doesn't know: MoreVoice may add more.

## Start a stream

### Stream a call's audio to your WebSocket

`POST /v1/calls/{id}/streams` · scope `calls:write` · [API reference](https://docs.morevoice.ai/api/operations/calls_create_stream/)

Starts a media stream on a live call in `fork` mode: a one-way copy of the caller (`inbound`), what the caller hears (`outbound`), or both, sent to your `wss://` URL as Twilio Media Streams messages (`connected`, `start`, `media`, `stop`). The call goes on exactly as before: if your server is slow, frames older than 2 s are dropped (`frames_dropped`); if it disconnects, the stream reconnects (3 tries) and otherwise gives up (`failed`) without touching the call. The upgrade request is signed with your signing secret over its path and query (`webhook-id` is the stream's ID). The stream stops when the call ends or with DELETE. At most 2 streams per call.

| Field | Required | What it does |
| --- | --- | --- |
| `custom_parameters` |  | Up to 20 string key/values the receiver gets in `start.customParameters` (as Twilio's stream parameters). |
| `format` |  | `mulaw_8000` (the default): G.711 μ-law at 8 kHz, Twilio's; `l16_16000` / `l16_8000`: signed 16-bit PCM, little-endian. |
| `mode` |  | `fork` (the default): a one-way copy; your server's messages are ignored. |
| `tracks` |  | `inbound`: the caller; `outbound`: what the caller hears (the assistant, an agent, prompts); `both` (the default). |
| `url` | Yes | Your WebSocket URL (`wss://`). The upgrade is signed with your signing secret (`webhook-id`, `webhook-timestamp`, `webhook-signature` over the path and query). |

**cURL**

```sh
curl -X POST https://api.morevoice.ai/v1/calls/call_7Hk2Lm9Qp/streams \
  -H "Authorization: Bearer $MOREVOICE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{
  "custom_parameters": {
    "crm_id": "42"
  },
  "format": "mulaw_8000",
  "tracks": "both",
  "url": "wss://media.example.com/streams"
}'
```

**Node.js**

```ts
import MoreVoice from "@morevoice/sdk";

const mv = new MoreVoice(); // MOREVOICE_API_KEY from the environment

const mediaStream = await mv.calls.createStream("call_7Hk2Lm9Qp", {
	custom_parameters: {
		crm_id: "42",
	},
	format: "mulaw_8000",
	tracks: "both",
	url: "wss://media.example.com/streams",
});
console.log(mediaStream);
```

**Python**

```python
from morevoice import MoreVoice

client = MoreVoice()  # MOREVOICE_API_KEY from the environment

media_stream = client.calls.create_stream("call_7Hk2Lm9Qp", {
    "custom_parameters": {
        "crm_id": "42",
    },
    "format": "mulaw_8000",
    "tracks": "both",
    "url": "wss://media.example.com/streams",
})
print(media_stream)
```

### Stop a media stream

`DELETE /v1/calls/{id}/streams/{stream_id}` · scope `calls:write` · [API reference](https://docs.morevoice.ai/api/operations/calls_delete_stream/)

Stops the stream: what is queued is flushed, your server gets `stop`, and the socket closes. The call is not affected. Streams also stop on their own when the call ends.

**cURL**

```sh
curl -X DELETE https://api.morevoice.ai/v1/calls/call_7Hk2Lm9Qp/streams/ms_4fG7hJ2kL9mN3pQ6rS8tUv \
  -H "Authorization: Bearer $MOREVOICE_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)"
```

**Node.js**

```ts
import MoreVoice from "@morevoice/sdk";

const mv = new MoreVoice(); // MOREVOICE_API_KEY from the environment

const mediaStream = await mv.calls.deleteStream("call_7Hk2Lm9Qp", "ms_4fG7hJ2kL9mN3pQ6rS8tUv");
console.log(mediaStream);
```

**Python**

```python
from morevoice import MoreVoice

client = MoreVoice()  # MOREVOICE_API_KEY from the environment

media_stream = client.calls.delete_stream("call_7Hk2Lm9Qp", "ms_4fG7hJ2kL9mN3pQ6rS8tUv")
print(media_stream)
```

> **Your bot, your responsibility**
>
> In `connect` mode your bot speaks to the caller in your name. Tell callers they are talking to an AI, honour opt-outs (send `hangup`, and add the number to the [do-not-call list](https://docs.morevoice.ai/guides/compliance/#honour-an-opt-out)), and follow the [Acceptable Use Policy](https://morevoice.ai/en/legal/aup/). MoreVoice still checks every outbound call before it is dialled.
