Documentation
Everything you need to build with VocaSync — text-to-speech synthesis, forced alignment, transcription, and translation. New here? Start with Getting started, or jump to a section below. Building an integration? See the API reference.
Introduction
Core services
Speech synthesis
Turn text into studio-quality speech — voices, quality tiers, output formats, pronunciation maps, and chunking.
Forced alignment
Deterministic word-level timestamps from your own transcript — JSON, SRT, and WebVTT output.
Transcription
Speech-to-text across 99+ languages, with transcript review before a linked alignment.
Translation
Neural machine translation with register-aware tone and subtitle readability limits.
Workflows
Subtitling
Transcription → alignment → translation in one bundle, with optional transcript review and upfront pricing.
Subtitle Base
Transcription → alignment in one bundle — timed cues in the language you recorded, no translation stage.
Quick Translate
Transcription → translation in one bundle — translated subtitles without forced alignment.
Dub-Prep
Synthesis → alignment in one bundle — generate narration and get word-level timings for dubbing.
Developers
API keys & authentication
Create keys, authenticate requests, and drive every VocaSync service from your own code.
Webhooks
Subscribe to job and pipeline events, filter what you receive, and verify signatures.
Publishable keys
Share artifacts publicly with scoped, revocable keys — no API secret exposure.
Integrations
Official plugins for Astro and WordPress, plus building your own integration.