# Alternatives to Fish Audio

Fish Audio is a text-to-speech and voice cloning platform built on the open-weight Fish Speech / OpenAudio models, cloning a voice from a short reference sample across 80+ languages. It targets developers and indie creators who want ElevenLabs-level quality with an API priced far below incumbents.

Fish Audio ranks #8 of 10 in AI voice generation, with an Alt Score of 85. It is available on the web. 8 of 9 checklist rows are verified against a public source.

Source: https://altcatalog.com/alternatives/fish-audio/
Category: AI voice generation

## Overview

- **Who it's for**: Developers and businesses building voice AI products—chatbots, virtual assistants, audiobooks, games, and localized/dubbed content—who need text-to-speech, voice cloning, and speech-to-text via an API or SDK, as well as individual creators who want to clone or generate voices through the web app.
- **What you get**: Text-to-speech models (S1, S2-Pro, S2.1-Pro) covering 80+ languages with inline emotion and paralinguistic control, instant or persistent voice cloning from short audio samples, speech-to-text transcription, a voice-design tool that generates candidate voices from text prompts, and WebSocket-based real-time streaming, all reachable through REST/WebSocket APIs, Python and JavaScript SDKs, or the browser-based app.
- **How it works**: Users submit text (or reference audio for cloning) to the Fish Audio API or web app; usage is billed pay-as-you-go per million UTF-8 bytes of input text for TTS and per audio hour for transcription, with no subscription fee. A voice can be cloned by uploading audio samples to train a reusable model that returns a voice ID for later text-to-speech calls, or used instantly by passing reference audio inline on a single request without training a persistent model.

## Profile

- **Pricing model**: Usage-based (verified 2026-07-30)
- **Starts at**: $15/M UTF-8 bytes (TTS) (verified 2026-07-30)
- **Platforms**: Web
- **Status**: active (verified 2026-07-30)

## Ranked alternatives

| # | App | Alt Score | Licence | Platforms |
|---|-----|-----------|---------|-----------|
| 1 | [Cartesia](https://altcatalog.com/alternatives/cartesia.md) | 100 | Proprietary | ? |
| 2 | [Speechify](https://altcatalog.com/alternatives/speechify.md) | 100 | Proprietary | iOS, Android, macOS, Windows |
| 3 | [Camb.ai](https://altcatalog.com/alternatives/camb-ai.md) | 95 | Proprietary | Web |
| 4 | [ElevenLabs](https://altcatalog.com/alternatives/elevenlabs.md) | 95 | Proprietary | Web, iOS, Android |
| 5 | [Murf](https://altcatalog.com/alternatives/murf.md) | 95 | Proprietary | Web |
| 6 | [Resemble AI](https://altcatalog.com/alternatives/resemble-ai.md) | 90 | MIT | Web, Browser extension |
| 7 | [Hume AI](https://altcatalog.com/alternatives/hume-ai.md) | 85 | Proprietary | Web |
| 8 | [WellSaid](https://altcatalog.com/alternatives/wellsaid.md) | 85 | Proprietary | Web |
| 9 | [Lovo (Genny)](https://altcatalog.com/alternatives/lovo.md) | 80 | ? | Web |

Alt Score = Verified coverage (90%) + Visibility (10%). See https://altcatalog.com/how-alt-score-works/

## Feature comparison

Legend: Yes / No / Partial / ? (not verified).

| AI voice generation checklist | Fish Audio | Cartesia | Speechify | Camb.ai | ElevenLabs | Murf |
|---|---|---|---|---|---|---|
| Pricing model | Usage-based | Freemium | Freemium | Freemium | Freemium | Freemium |
| Starts at | $15/M UTF-8 bytes (TTS) | $5/mo | $29/month | $5/mo | $6/month | ? |
| License | ? | Proprietary | Proprietary | Proprietary | Proprietary | Proprietary |
| Platforms | Web | ? | iOS, Android, macOS, Windows, Web, Browser extension | Web | Web, iOS, Android | Web |
| Voice cloning | Yes | Yes | Yes | Yes | Yes | Yes |
| Languages & accents | Yes | Yes | Yes | Yes | Yes | Yes |
| Emotional control | Yes | Yes | Yes | Yes | Yes | Yes |
| API access | Yes | Yes | Yes | Yes | Yes | Yes |
| Commercial license terms | ? | Yes | Yes | Partial | Partial | Yes |
| Free tier limits | Yes | Yes | Yes | Yes | Yes | Yes |
| Real-time / streaming | Yes | Yes | Yes | Yes | Yes | Partial |
| Dubbing (video translation) | Partial | Yes | Yes | Yes | Yes | Yes |
| Ethics & consent policy | Yes | Yes | Yes | Yes | Yes | Yes |

## Sources

Sources for Fish Audio. Each alternative is sourced on its own page.

- **Pricing model**: Usage-based — <https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md> (verified 2026-07-30)
  - Quote: “The Fish Audio API uses pay-as-you-go pricing based on actual usage. There are no subscription fees or monthly minimums for API access.”
- **Starts at**: $15/M UTF-8 bytes (TTS) — <https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md> (verified 2026-07-30)
  - Note: Pay-as-you-go, no subscription; a free dev model (s2.1-pro-free) is $0 under fair-use limits. No fixed monthly plan exists.
  - Quote: “s2.1-pro | $15.00 / M UTF-8 bytes”
- **Platforms**: Web — <https://docs.fish.audio/developer-guide/core-features/creating-models.md> (verified 2026-07-30)
  - Note: Also accessible via REST/WebSocket API and Python/JavaScript SDKs; no dedicated desktop or mobile apps found.
  - Quote: “Use it in the web app ... No code — clone a voice in the browser.”
- **Status**: active — <https://fish.audio/sitemap.xml> (verified 2026-07-30)
  - Note: Sitemap last-modified matches current date; docs also list a live changelog and a newest model (s2.1-pro) as the current recommended production model.
  - Quote: “<lastmod>2026-07-30T04:09:18.144Z</lastmod>”
- **Voice cloning**: Yes — <https://docs.fish.audio/developer-guide/core-features/creating-models.md> (verified 2026-07-30)
  - Quote: “Build a reusable voice model from your own audio, then use it anywhere you generate speech. You get back a voice id — pass it as reference_id to Text to Speech.”
- **Languages & accents**: Yes — <https://docs.fish.audio/developer-guide/models-pricing/models-overview.md> (verified 2026-07-30)
  - Quote: “S2.1-Pro supports 83 languages, while S2-Pro supports 80+ languages. Both use automatic language detection and support inline emotion and paralinguistic cues.”
- **Emotional control**: Yes — <https://docs.fish.audio/developer-guide/models-pricing/models-overview.md> (verified 2026-07-30)
  - Quote: “S2.1-Pro and S2-Pro treat [bracket] tags as standard text rather than dedicated control tokens... you can use any descriptive expression and the model will interpret it, such as [whispers sweetly] or ”
- **API access**: Yes — <https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md> (verified 2026-07-30)
  - Note: Full REST + WebSocket API with Python and JavaScript SDKs is documented at docs.fish.audio/api-reference.
  - Quote: “The Fish Audio API uses pay-as-you-go pricing based on actual usage.”
- **Free tier limits**: Yes — <https://docs.fish.audio/developer-guide/models-pricing/models-overview.md> (verified 2026-07-30)
  - Note: Free tier is scoped to dev/testing/smaller businesses, not full production SLA.
  - Quote: “Fish Audio S2.1-Pro Free - The same model as S2.1-Pro, available at $0 for development and testing ... Free to use under fair-use limits ... No TTFA or DPA guarantees.”
- **Real-time / streaming**: Yes — <https://docs.fish.audio/developer-guide/best-practices/real-time-streaming.md> (verified 2026-07-30)
  - Quote: “Real-time streaming lets you generate speech as you type or speak, perfect for chatbots, virtual assistants, and live applications.”
- **Dubbing (video translation)**: Partial — <https://docs.fish.audio/developer-guide/core-features/creating-models.md> (verified 2026-07-30)
  - Note: Voice cloning supports cross-language dubbing use cases via the API, but no dedicated video-upload/auto-dub product feature was found in the fetched docs.
  - Quote: “Dubbing & localization: Keep a speaker's identity across languages.”
- **Ethics & consent policy**: Yes — <https://docs.fish.audio/developer-guide/best-practices/voice-cloning.md> (verified 2026-07-30)
  - Quote: “Only clone voices you have permission to use: Your own voice; Someone who gave you written permission; Never use voices from the internet without permission; Never use celebrity or public figure voice”

---
Ranked by verified data, never by who paid. https://altcatalog.com/trust/