The model catalog
Assistants don’t hardcode providers — they pick from a curated model catalog covering four types:
Only the platform admin controls availability. They enable models globally,
can grant or block a model for one user, and define which engine modes each model
supports. Tenant admins cannot override the catalog.
With the Fallbacks & Guardrails add-on, the assistant editor shows provider and model selectors.
Pipeline, realtime, and half-cascade keep independent choices. Choosing the
default stores
null, so later admin changes take effect automatically:
- Assistant selection (Fallbacks & Guardrails only)
- Personal default set by the platform admin
- Platform default
- Safe runtime fallback
The voice library
The voice picker loads the complete paginated catalog from every enabled TTS provider, filterable by language and gender, with instant audio previews — click play to hear a sample before committing. ElevenLabs voices with custom rates or credit multipliers are excluded, so every ElevenLabs voice shown in the picker uses the provider’s standard voice rate.- Voice per assistant — the default speaking voice.
- Voice override per flow agent — individual flow agents can speak with a different voice (e.g. a distinct voice for the “sales specialist” agent).
- Voice per language — with automatic language switching, you map a voice per language so the assistant switches voice together with language.
- Cloned voices — with the Clone your own voice add-on (or Include free on the plan), you can add custom clones and use them like any library voice. A clone is private to the account that created it: other members of the same workspace cannot see or select it. Capacity is limited by
max_cloned_voices.
Speaking style
Pipeline and half-cascade assistants use the same provider-aware voice controls. The editor only sends settings supported by the resolved TTS model:- Cartesia Sonic 3 / 3.5 — speaking rate, output volume, and optional dynamic emotion cues.
- Eleven v3 — stability and native audio-tag expressiveness. Unsupported speed, similarity, style, and speaker-boost controls are not sent.
- Other supported ElevenLabs models — speaking rate, stability, similarity, style exaggeration, and speaker boost.
- OpenAI and Gemini TTS — free-text speaking-style instructions when the selected model supports them.
- Pronunciation dictionary — per-assistant key-value mapping, e.g.
"API" → "A P I"or your company name spelled phonetically. Applied to all spoken output. - Markdown & emoji filters — enabled by default so LLM formatting artifacts are never read aloud.
Fallback chains (Fallbacks & Guardrails)
For production reliability you can define a fallback chain per stage: if the primary provider degrades mid-call, the engine switches to the next entry — the call continues seamlessly.- STT — e.g.
Deepgram nova-3 → AssemblyAI universal-streaming - LLM — e.g.
GPT → Groq - TTS — e.g.
Cartesia sonic → ElevenLabs(each entry may specify its own voice)
Even without configured fallbacks, a broken configuration never kills a call: the engine falls back to a safe default stack and records a
config_error event.