Pika Audio Models: Soundtrack, Music, SFX, and Speech
Pika Labs
Pika Labs shipped four foundation audio models in its Pika Audio family. Pika Speech is a 3B flow-matching transformer producing 48 kHz studio-quality speech with seconds-of-reference voice cloning at a 0.02 real-time factor; Pika SFX turns text into focused sound effects in real time; Pika Music generates finished tracks from text, lyrics, or reference audio; and Pika Soundtrack turns video into motion-aware synchronized soundscapes. The family is positioned up to 20x cheaper than comparable audio models and is available via the Pika API Club.
Why it matters
Pika — a video-first lab — is making a full-stack play across TTS, music, SFX, and soundtracks at a fraction of incumbent pricing, which directly pressures ElevenLabs, Suno, and Stability Audio.
Importance: 3/5
Full-stack audio family at incumbent pricing pressure