RodiumAi docs
Core concepts

Text to speech

Send JSON with model and input; receive audio bytes. Billed in RODI (incl. output_audio when priced).

POSThttps://api.rodiumai.io/v1/audio/speech

Call POST /v1/audio/speech with a TTS model id. OpenAI voices (alloy, …) or Gemini prebuilt voices (Kore, Puck, …).

When to use

  • IVR prompts and product narration in French or English.
  • Choose Gemini expressive voices vs OpenAI tts-1 for cost.
  • Save the binary response directly to a file.

Recipes

French narration

input in French, voice Kore (Gemini) or alloy (OpenAI), response_format wav or mp3.

Note: Write the response body as binary — do not JSON.parse it.

Gemini vs OpenAI

google/gemini-*-tts* for Vertex/Google; openai/tts-1 when you need classic OpenAI voices.

Note: Confirm the slug in GET /v1/models before shipping.

WAV note (Gemini)

Gemini upstream is PCM; Rodium wraps WAV for non-pcm formats so browsers can play the file.

Note: Use response_format=pcm only if you decode L16 yourself.

Examples

Gemini TTS

OpenAI TTS

Request parameters

ParameterTypeRequiredDescription
modelstringRequiredTTS model id (e.g. openai/tts-1, google/gemini-2.5-flash-preview-tts).
inputstringRequiredText to synthesize into speech.
voicestringOptionalVoice id when supported (e.g. "alloy", "nova").
response_formatstringOptionalAudio format when supported (e.g. "mp3", "opus", "wav").
speednumberOptionalPlayback speed multiplier when supported (typically 0.25–4.0).
instructionsstringOptionalOptional speaking style / delivery instructions when the model supports them.

API reference: speech