Aura relies on state-of-the-art Text-to-Speech (TTS) models to deliver human-like, low-latency audio.
Supported Providers
Seminara uses a multi-provider architecture for maximum reliability and lowest latency:
- Primary Voice Provider: Optimized for ultra-low latency conversational agents. This is the default provider.
- Fallback Voice Provider: Used if the primary provider experiences downtime or if a highly specific emotive voice is required.
Selecting a Voice
In the Aura Customization settings, you can preview the available voices. We provide a curated list of voices optimized for presentations. Currently available voices:
- Female 1 (Thalia) - Female, US
- Male 1 (Orion) - Male, US
Latency Considerations
Our orchestration engine is tuned to achieve 1200ms - 2000ms response times. Voice selection can slightly impact this. Our primary provider generally offers the lowest time-to-first-byte (TTFB), which is critical for making interruptions feel natural.
Frequently Asked Questions (FAQ)
Q: Can I clone my own voice?
A: Voice cloning is currently available exclusively on the Ultra tier. Contact support to enable this feature.
Q: Does the voice support different accents?
A: Yes, while the default voices are US English, you can explicitly instruct Aura in the Custom Instructions to adopt a British or Australian accent, and the generation model will adapt accordingly.
Q: Can I change the speed of the voice?
A: While hosts cannot set a global speed multiplier, attendees can verbally ask Aura to "speak a bit slower" or "speed up" during the session, and she will dynamically adjust her pacing.
Q: What if the voice mispronounces my company name?
A: You can provide phonetic spellings in your Custom Instructions (e.g., "Pronounce Seminara as Sem-in-ARR-ah"). Aura will strictly follow this phonetic guide when speaking.