We ship open weights. That means our models can be fine-tuned on arbitrary speech data, and voice cloning is within the technical capability of anyone who downloads them. We need to be direct about the risks this creates and explicit about the framework we use to govern it. This post explains our principles and, more importantly, the reasoning behind them.
What voice cloning actually requires with our open weights
Fine-tuning our acoustic model on a target speaker's voice requires roughly 3-8 hours of clean, transcribed audio from that speaker. This is not a trivial threshold. It limits casual misuse but does not prevent determined bad actors with access to recordings of a specific individual. For public figures who speak on camera regularly, 3-8 hours of audio is potentially accessible. For private individuals, it is substantially harder to obtain.
The vocoder component can be fine-tuned on less data (sometimes under an hour) to adapt timbre more specifically. This means voice cloning in practice has a spectrum: rough voice approximation (achievable with less data, less convincing) versus high-quality reproduction (requiring more data and more training). Neither category is acceptable without consent, but they carry different practical risk levels.
Our consent framework
For voice cloning features accessible through our API (as opposed to the open weights path, which we cannot control post-download), we require explicit consent from the speaker whose voice is being cloned. This means:
The voice enrollment flow requires the speaker to record a specific consent phrase that is then verified before any cloning can proceed. The consent recording is retained and associated with the voice profile. Third parties cannot enroll a voice on behalf of the speaker through our API; only the speaker themselves can initiate enrollment.
Voice profiles created through our API can only be used by the account that created them, and only for the use case the account has registered (we ask for use-case declaration during onboarding). We do not allow voice profiles to be transferred or sold through our platform.
This framework is not technically unbreakable. A sophisticated actor could construct a consent workflow that obtains a recording of someone else's consent phrase. We are not claiming our controls prevent all misuse. We are claiming they raise the barrier enough to eliminate casual misuse through our platform while making deliberate misuse require active circumvention.
The open weights problem
Here is the honest tension: we ship open weights specifically because we believe open models are better for the ecosystem, enable legitimate use cases that a closed API cannot serve, and allow researchers and developers to work with our models in ways we cannot anticipate. We are not going to walk back the open-weights commitment.
But open weights mean we cannot enforce our consent framework on fine-tuning that happens outside our platform. Anyone with a GPU and our model weights can fine-tune on any audio data they have access to. We accept this as a consequence of the open-weights decision.
What we can do: be explicit that cloning a person's voice without their consent is a harm we actively oppose, document our consent framework clearly so it is a model for implementers building on our weights, and decline to provide technical assistance or tooling that specifically facilitates non-consensual cloning. We also plan to contribute watermarking research that makes it possible to identify audio generated by fine-tuned versions of our models, though this is not complete work and we will not claim it as a current safeguard.
Why platform-level controls matter more than model-level restrictions
There is a recurring argument in the voice AI space that models should be technically restricted to prevent cloning: limited fine-tuning capability, minimum data requirements enforced at inference time, speaker embeddings that can only be extracted with specific permissions. We have thought about this carefully and concluded that model-level restrictions are insufficient and have significant costs.
Model restrictions can be removed or circumvented by anyone working with open weights directly. The bad actors who would misuse voice cloning have the technical capability to work around model-level restrictions. The people who are harmed by model-level restrictions are legitimate users with uncommon but valid use cases that the restrictions do not accommodate.
Platform-level controls, combined with legal frameworks (many jurisdictions are developing voice impersonation and deepfake regulations), social norms around consent, and technical detection tools, are a more effective combined defense than model restrictions alone. We think the AI voice space needs more investment in detection and attribution tooling and less investment in capability restriction that harms legitimate use.
What we are watching
India does not currently have specific legislation addressing voice deepfakes or non-consensual AI voice cloning. The Information Technology Act and its amendments provide some broad coverage, but the application to voice-specific harms is untested. We expect this to change as AI-generated voice becomes more prevalent in fraud, political disinformation, and harassment cases. We are engaging with legal advisors on how our terms of service and consent framework should evolve as the legal landscape develops.
For Arabic voice specifically: several Gulf states are developing AI governance frameworks with stronger provisions around impersonation. We are monitoring those developments and expect our consent framework to need adjustment as those regulations firm up.