After registering and opening Voice Clone, the page asks you to record 25 phrases in sequence.
The collection process does not take long.
After I finished recording and submitted the samples, the voice model completed training in approximately 30 minutes.
When I began generating speech, the most obvious issue was language support.
I initially wanted to test Chinese directly, but the web version displayed a message saying it was unsupported. After looking into it, I found that MyVocal originally focused primarily on English, with additional languages gradually introduced through paid versions and later models.
The current V3 API lists 98 supported languages, including Chinese. However, whether the web product provides access to the same model cannot be determined from the API documentation alone.
The English test went more smoothly.
I entered an ordinary piece of text. The generated voice had a tone that was recognizably similar to the original, and the overall delivery was relatively smooth.
However, as the sentences became longer, the AI quality became more obvious.
Pauses, emphasis, and emotional changes were less natural than a real human recording. This was particularly noticeable in sentences requiring significant emotional variation. The voice sounded more like it was reading the text fluently than having a normal conversation.
This means that “does the voice sound like you?” and “does it sound natural?” are actually separate questions.
MyVocal can achieve the former to some extent, but the latter still reveals many signs of synthesis.
Privacy deserves more attention than generation speed.
MyVocal previously faced media questions about the transparency of its operations. Its operating entity can now be identified, but voice cloning still requires users to provide a set of recordings containing personally identifiable vocal characteristics.
For casual entertainment testing, evaluating the risk may be relatively straightforward.
For work recordings, customer information, or long-term commercial use, you should first confirm how recordings are stored, how deletion works, and what rights the service receives before uploading anything.
Pros
- A short cloning process: Training can begin after recording only 25 phrases, without requiring a large audio dataset.
- AI Singer is distinctive: In addition to ordinary text-to-speech, the cloned voice can be used for singing.
- The free plan lets you test the complete workflow: Useful for evaluating voice similarity before paying.
- Provides an API: Developers can integrate voice-generation capabilities into their own products.
- Low barrier to entry: No professional recording skills or knowledge of voice models is required.
Cons
- Chinese support on the web version is unclear: The API has expanded its language coverage, but web-product functionality needs to be verified separately.
- Limited naturalness: Long sentences, pauses, and emotional delivery can still sound synthetic.
- Historical pricing is confusing: Prices vary considerably between different periods and sources.
- Voice data raises privacy concerns: Users should carefully review data handling and licensing terms before using cloning features.
- Limited professional controls: Control over speed, pauses, and detailed delivery is less advanced than on specialist voice platforms.
Best for / Not ideal for
Best for
- People who want to try voice cloning: Useful for seeing how closely AI can imitate their voice.
- People interested in AI singing: AI Singer provides a straightforward entertainment feature.
- Creators who need a quick voice demo: A cloned voice can be used to test narration and content pacing.
- Developers who need a voice API: Suitable for quickly adding voice cloning or text-to-speech to a product.
Not ideal for
- People primarily producing professional Chinese voiceovers: Chinese support and practical output quality need to be tested separately.
- People with strict voice-data security requirements: The platform’s data-processing practices should be confirmed before uploading a personal voice.
- People who need broadcast-quality output: Naturalness and emotional control still cannot fully replace a human voice actor.
- Teams requiring fine-grained voice control: MyVocal may not offer enough control when every sentence needs adjustments to pauses, speed, and intonation.
Comments (0)