Pika launches Pika Audio: four frontier foundation sound models at up to 20x lower cost
Pika
On August 14, 2026, Pika released Pika Audio, its first family of frontier foundation sound models, comprising Soundtrack (video-to-audio, $0.617/sec), Music (up to 6-minute songs from text/lyrics/voice/reference), SFX (text-to-sound-effects, 44.1 kHz stereo, 0.847s avg latency), and Speech (48 kHz TTS with preset voices or voice cloning, real-time factor 0.02). Pika markets the family as up to 20x cheaper than comparable audio models, with specific claims of 9x vs ElevenLabs v3, 4.5x vs Cartesia/ElevenLabs Turbo, and 2x vs Hunyuan Foley / Fish Audio.
Why it matters
Marks Pika's expansion from video-only into a full audio foundation model suite, with the Speech model pitched as the cheapest credible ElevenLabs competitor on the market and the Soundtrack model positioned as a cost-efficient alternative to HunyuanVideo-Foley for video-to-audio. Gated behind the new Pika API Club ($10/month membership) that aggregates 70-100+ third-party models.
Importance: 3/5
4 independent confirmations; new audio foundation family (default base 2)