Tips for creating your custom voice September 09, 2026 14:56 Updated You can now create your very own voice clones. This feature is available for Enterprise and Agency users only. Check the following tips and tricks to produce the best results. Preparing your audio samples for upload Reading your consent statement Submission Preparing your audio samples for uploadBefore you can create your own custom voice, you must upload some audio samples for the model to train and generate the custom voice. The following tips can help create a more authentic and engaging digital representation for your custom voice. Do’s Record 1 to 2 minutes of clear audio Record in a quiet, echo-free environment Stick to one language Keep tone natural and pace steady Use a mix of sentence styles such as questions and statements Don’ts Don’t record more than 3 minutes Don't record with in an area with background noise and music Don’t whisper or shout Don’t switch language or dialects Here are some best practices to help you record your audio samples to create the best possible voice clone:Record at least 1 minute of audioAvoid recording more than 3 minutes, this will yield little improvement and can, in some cases, even be detrimental to the clone.How the audio was recorded is more important than the total length (total runtime) of the samples. The number of samples you use doesn’t matter; it is the total combined length (total runtime) that is the important part.Approximately 1-2 minutes of clear audio without any reverb, artifacts, or background noise of any kind is recommended. When we speak of “audio or recording quality,” we do not mean the codec, such as MP3 or WAV; we mean how the audio was captured. However, regarding audio codecs, using MP3 at 128 kbps and above is advised. Higher bitrates don’t have a significant impact on the quality of the clone.Keep the audio consistentThe AI will attempt to mimic everything it hears in the audio. This includes the speed of the person talking, the inflections, the accent, tonality, breathing pattern and strength, as well as noise and mouth clicks. Even noise and artifacts which can confuse it are factored in.Ensure that the voice maintains a consistent tone throughout, with a consistent performance. Also, make sure that the audio quality of the voice remains consistent across all the samples. Even if you only use a single sample, ensure that it remains consistent throughout the full sample. Feeding the AI audio that is very dynamic, meaning wide fluctuations in pitch and volume, will yield less predictable results.Replicate your performanceAnother important thing to keep in mind is that the AI will try to replicate the performance of the voice you provide. If you talk in a slow, monotone voice without much emotion, that is what the AI will mimic. On the other hand, if you talk quickly with much emotion, that is what the AI will try to replicate.It is crucial that the voice remains consistent throughout all the samples, not only in tone but also in performance. If there is too much variance, it might confuse the AI, leading to more varied output between generations.Find a good balance for the volumeFind a good balance for the volume so the audio is neither too quiet nor too loud. The ideal would be between -23 dB and -18 dB RMS with a true peak of -3 dB.File formatsMake sure you record your audio samples in the following formats: MP3 WAV M4A Make sure your files are no longer than 15MB each and that the total durations of all of the audio samples do not exceed 3 minutes long.Reading your consent statementAfter reviewing your audio samples, you must use your device’s mic to make a live recording of the following consent statement:“My name is [full name]. By recording this audio, I give Vyond permission to use my voice samples to create a digital copy of my voice, which only I will have access to. The passcode is [passcode]” Replace [FULL NAME] with your name. The statement must be clearly audible and understandable. We currently only accept consent statements in ENGLISH. Read out the randomly generated pass code clearly and accurately. Your voice clone generation cannot be processed if the required consent statement is not read out correctly, or the consent checkbox is not checked. Once our system has verified that you have read out the consent statement correctly, you will be able to proceed to the submission stage to generate your voice clone.SubmissionOnce the audio files are and the consent recording have been submitted, you can generate your voice clone. The voice cloning process takes less than 20 seconds. Related articles How do I create my own voice? Known Issue - “Latest” Voices Failing to Generate Audio (RESOLVED)