Skip to main content

Quick Start

Convert text into natural-sounding speech with customizable voices and emotions.
1

Navigate to New Voice

Navigate to the new voice generator page.
new-voice
2

Select Model

Choose the AI model you want to use for your voice generation
Select Model
3

Write Your Text

Enter the text or script you want to hear in the voice-over section.
Write Your Text
4

Select Voice Over Template

Browse the sample templates and choose the voice that best fits your content’s tone and style.
Select Voice Over Template
5

Click Generate

Click the Generate button to process your text into the selected voice.
Click Generate
6

Download Your Audio

Once processing is complete, download and save your professional AI voiceover.

Voice Selection

Voice Types

Characteristics:
  • Deep, authoritative tones
  • Professional narration
  • Various age ranges
  • Multiple accents
Best For:
  • Corporate narration
  • Audiobooks
  • Professional content
  • Educational material

Language Support

Zoice supports 50+ languages including:
  • English: US, UK, Australian, Indian, and more
  • Spanish: Spain, Latin American variants
  • French: France, Canadian
  • German: Standard, Austrian, Swiss
  • Mandarin: Simplified, Traditional
  • Japanese, Korean, Hindi, Arabic, and many more

Voice Customization

Speed Control

  • Slow (0.5x-0.8x): Deliberate, clear enunciation
  • Normal (1.0x): Natural speaking pace
  • Fast (1.2x-1.5x): Quick, energetic delivery

Pitch Adjustment

  • Lower (-20% to -10%): Deeper, more authoritative
  • Normal (0%): Natural voice pitch
  • Higher (+10% to +20%): Lighter, more energetic

Emotion & Tone

Neutral

Standard, professional delivery

Happy

Upbeat, cheerful, positive

Sad

Somber, melancholic, gentle

Excited

Energetic, enthusiastic, dynamic

Calm

Soothing, relaxed, peaceful

Serious

Professional, formal, authoritative

Text Formatting

Punctuation Control

Punctuation affects pacing and delivery:

Special Formatting

Use asterisks or capitals for emphasis:
Add pauses with special tags:
Guide pronunciation with phonetic spelling:
Format numbers for proper reading:

Use Cases

Voiceovers & Narration

Customer Service

E-Learning

Audiobooks

Best Practices

  • Write conversationally, not formally
  • Use shorter sentences
  • Avoid complex punctuation
  • Read aloud to test flow
  • Use natural language
  • Spell out acronyms first time
  • Provide pronunciation guides for unique terms
  • Avoid ambiguous abbreviations
  • Use standard formatting for numbers
  • Use punctuation for natural pauses
  • Break long paragraphs into shorter segments
  • Vary sentence length
  • Add pauses for emphasis
  • Match voice to content tone
  • Consider target audience
  • Test different voices
  • Stay consistent within projects

Advanced Features

SSML Support

Use Speech Synthesis Markup Language for advanced control:

Multi-Voice Scripts

Use different voices in one audio:

Background Music

Add background music to your voice:
  • Upload music track
  • Set volume levels
  • Auto-duck during speech
  • Fade in/out options

Audio Output Settings

Format Options

  • MP3: Compressed, smaller file size
  • WAV: Uncompressed, highest quality
  • OGG: Open format, good compression

Quality Settings

  • Standard: 128 kbps (web, streaming)
  • High: 192 kbps (general use)
  • Premium: 320 kbps (professional)

Sample Rate

  • 22 kHz: Voice-only, smaller files
  • 44.1 kHz: CD quality, standard
  • 48 kHz: Professional, broadcast

Common Scenarios

YouTube Videos

Narration for video content

Podcasts

Intro/outro or full episodes

IVR Systems

Phone menu and prompts

Accessibility

Make content accessible

Troubleshooting

  • Add phonetic spelling
  • Break compound words
  • Use pronunciation guides
  • Try different voice
  • Adjust speed setting
  • Add punctuation for pauses
  • Break long sentences
  • Use pause tags
  • Add emotion settings
  • Use varied punctuation
  • Try different voice
  • Add emphasis markers
  • Use emphasis markers
  • Adjust punctuation
  • Rewrite sentence structure
  • Use SSML tags
Always preview a short sample before generating long audio files to ensure the voice and settings are correct.

Character Limits

  • Single Generation: Up to 5,000 characters
  • Batch Processing: Up to 50,000 characters
  • Audiobook Mode: Unlimited with chapter breaks

Next Steps

AI Avatar

Combine voice with avatars

Prompt Guide

Master voice prompt engineering