# AI Voice

Clone a voice from a recording, then use it to read any text out loud — and read your customers back, turning what they said into text your agent can answer.

Three resources cover the whole audio loop. An **AI Voice** is the cloned voice itself. An **AI Speech** is one text read out loud by that voice. An **AI Transcript** is the text of an audio file you upload, from any speaker, cloned or not.

Everything travels as base64 inside the JSON body — no multipart uploads, no signed URLs to juggle. The audio format is read from the file's own header, so you never declare it.

Attach a voice to an [AI Agent](https://docs.starkinfra.com/get-started/ai-agent.md) and the loop closes: transcribe the caller, post the text, and synthesize the reply's `speech` field, which the agent already wrote to be heard rather than read.

NOTE: Read [Core Concepts](https://docs.starkinfra.com/get-started/core-concepts.md) before continuing this guide.

**RESOURCE SUMMARY**

**AI Voice**A voice cloned from a recording you upload.**AI Speech**Text read out loud by one of your voices, returned as a base64 MP3.**AI Transcript**The text of an audio file you upload. No voice required.

## Available Languages

Code samples on this page are in Python. The same content is available with samples in:

- [Python](https://docs.starkinfra.com/get-started/ai-voice-python.md)
- [Node.js](https://docs.starkinfra.com/get-started/ai-voice-node.md)
- [PHP](https://docs.starkinfra.com/get-started/ai-voice-php.md)
- [Java](https://docs.starkinfra.com/get-started/ai-voice-java.md)
- [Ruby](https://docs.starkinfra.com/get-started/ai-voice-ruby.md)
- [Elixir](https://docs.starkinfra.com/get-started/ai-voice-elixir.md)
- [.NET](https://docs.starkinfra.com/get-started/ai-voice-dotnet.md)
- [Go](https://docs.starkinfra.com/get-started/ai-voice-go.md)
- [Clojure](https://docs.starkinfra.com/get-started/ai-voice-clojure.md)
- [cURL](https://docs.starkinfra.com/get-started/ai-voice-curl.md)

## Setup

For each environment (Sandbox or Production):

1. Create a workspace at Stark Infra and generate your ECDSA keys.

2. Get in touch with your account manager to enable the AI products on your workspace.

3. Record a clean sample of the speaker talking naturally — no background music, no crosstalk. The clone is only as good as the recording.

## Typical flow

**1.** Upload the recording as an AI Voice. The call returns immediately with `status=processing`.

**2.** Poll `GET /v2/ai-voice` until the voice reaches `success`. Cloning usually takes seconds.

**3.** Send text and the voice id to `POST /v2/ai-speech`. The response carries the audio as a base64 MP3, synchronously.

**4.** Going the other way, send an audio file to `POST /v2/ai-transcript` and get its text back in the same call.

**5.** To give an agent a voice, set the voice id as the agent's `voiceId` and synthesize each reply's `speech` field.

## Audio encoding

Every audio payload — the recording you upload, the file you transcribe, the speech you get back — is a base64 string inside the JSON body.

The container is detected from the audio's own header. MP3, WAV, OGG, FLAC and WebM are recognized; anything else is treated as WebM. You never send a content type for the audio itself.

A payload is capped at 10000000 base64 characters, roughly 7.5 MB of audio. Synthesized speech always comes back as MP3.

## Use cases

**Voice support line:** transcribe the caller, run the text through an agent, and answer in your brand's own cloned voice.

**Audio notifications:** read balances, confirmations or alerts out loud in the same voice your customers already know.

**Call intake:** transcribe voice messages so your team reads instead of listens, and so your systems can search them.

**Accessibility:** offer a spoken version of any text your product already shows.

## AI Voice Overview

Here we show you how to clone a voice from a recording and how to know when it is ready to speak.

### Cloning a voice

`POST /v2/ai-voice`

Only `audio` is mandatory: a base64 recording of the speaker. Everything else — `name`, `description`, `language`, `gender` — is bookkeeping that makes the voice easy to find later.

`language` defaults to `portuguese`; send `english` for an English speaker. The call returns immediately with `status=processing`.

**Request**

```python
import base64
import starkinfra

with open("helena.wav", "rb") as recording:
    audio = base64.b64encode(recording.read()).decode()

voice = starkinfra.aivoice.create(
    starkinfra.AiVoice(
        name="Helena",
        description="Calm Brazilian Portuguese voice for customer support",
        language="portuguese",
        gender="female",
        audio=audio
    )
)

print(voice)
```

**Response**

```python
AiVoice(
    audio=None,
    created=2022-01-01 00:00:00,
    description=Calm Brazilian Portuguese voice for customer support,
    errors=[],
    gender=female,
    id=5656565656565656,
    language=portuguese,
    name=Helena,
    status=processing,
    updated=2022-01-01 00:00:00
)
```

### Waiting for the clone

`GET /v2/ai-voice`

Poll the list until the voice leaves `processing`. There is no webhook for this transition, and cloning usually takes seconds.

A voice in `failed` status carries the reason in `errors`. Only a voice in `success` can be used for speech — sending a processing voice to `POST /v2/ai-speech` is rejected.

**Request**

```python
import starkinfra

voices = starkinfra.aivoice.query()

for voice in voices:
    print(voice)
```

**Response**

```python
AiVoice(
    audio=None,
    created=2022-01-01 00:00:00,
    description=Calm Brazilian Portuguese voice for customer support,
    errors=[],
    gender=female,
    id=5656565656565656,
    language=portuguese,
    name=Helena,
    status=success,
    updated=2022-01-01 00:00:12
)
```

## AI Speech Overview

Speech is text-to-speech: you send the text and a voice id, and the same call returns the audio. Nothing to poll.

### Reading text out loud

`POST /v2/ai-speech`

The response carries the synthesized audio in `audio`, as a base64 MP3, alongside the speech record itself.

When the text comes from an agent, send the reply's `speech` field rather than its `text`. The agent writes `speech` to be heard: no Markdown, no URLs, no code, at most two sentences. Feeding it raw Markdown is the fastest way to make a good voice sound wrong.

**Request**

```python
import starkinfra

speech = starkinfra.aispeech.create(
    starkinfra.AiSpeech(
        voice_id="5656565656565656",
        text="I am sorry about the delay. I can open a delivery check for order 1234 right now."
    )
)

print(speech)
```

**Response**

```python
AiSpeech(
    audio=SUQzBAAAAAAAI1RTU0UAAAAPAAADTGF2ZjU4Ljc2LjEwMAAAAAAAAAAAAAAA...,
    created=2022-01-01 00:00:00,
    errors=[],
    id=1212121212121212,
    status=success,
    text=I am sorry about the delay. I can open a delivery check for order 1234 right now.,
    updated=2022-01-01 00:00:01,
    voice_id=5656565656565656,
    voice_name=None
)
```

### Fetching the audio again

`GET /v2/ai-speech/:id`

Every speech is stored, so you can download its audio again by id without paying for a second synthesis.

Send `fields` without `audio` when you only want the record — that keeps the file out of the response.

**Request**

```python
import starkinfra

speech = starkinfra.aispeech.get("1212121212121212")

print(speech)
```

**Response**

```python
AiSpeech(
    audio=SUQzBAAAAAAAI1RTU0UAAAAPAAADTGF2ZjU4Ljc2LjEwMAAAAAAAAAAAAAAA...,
    created=2022-01-01 00:00:00,
    errors=[],
    id=1212121212121212,
    status=success,
    text=I am sorry about the delay. I can open a delivery check for order 1234 right now.,
    updated=2022-01-01 00:00:01,
    voice_id=5656565656565656,
    voice_name=None
)
```

## AI Transcript Overview

Transcript is speech-to-text, the other half of the loop. No voice is involved: it reads any speaker, not only the ones you cloned.

### Transcribing audio

`POST /v2/ai-transcript`

Send the audio as base64 and get its text back in the same call. The format is detected from the file's own header.

This is step one of a voice channel: transcribe what the caller said, post that text to an AI Chat, and synthesize the reply's `speech` with one of your voices.

**Request**

```python
import base64
import starkinfra

with open("call.wav", "rb") as recording:
    audio = base64.b64encode(recording.read()).decode()

transcript = starkinfra.aitranscript.create(
    starkinfra.AiTranscript(
        audio=audio
    )
)

print(transcript)
```

**Response**

```python
AiTranscript(
    audio=None,
    created=2022-01-01 00:00:00,
    errors=[],
    id=9898989898989898,
    status=success,
    text=My order 1234 was supposed to arrive yesterday and it did not. What can I do?,
    updated=2022-01-01 00:00:02
)
```
