Most AI video gives you a different stranger every time. Persona Studio builds you a single presenter — a face you choose, a model trained on that face, and a voice that stays put — and then reuses it across your talking-head videos.
Ask us about availability on your plan.
A stock avatar changes every render, so nothing compounds. A trained presenter is the same person in video 1 and video 60 — which is what turns a feed into a brand.
The same trained face and the same voice in every talking-head video you make.
Start from an archetype, then set gender, age range, look, voice style, energy and language.
An unobtrusive "AI generated" label ships on the videos that feature your presenter.
You make three choices. Everything between them is automatic, and nothing renders before you've approved the face.
Start from 52 ready-made presenter archetypes across 7 industry families — local trades, health & beauty, professional services, food & hospitality, retail & lifestyle, content niches and affiliate. Then tune six things: gender, age range, look, voice style, energy and language.
We generate four photorealistic candidate faces from your archetype. You pick the one that feels like your brand — nothing is trained until you do.
Your chosen face is expanded into a training set of 10–20 pose and expression variants, then used to train a model of that specific face. It usually takes about 5–10 minutes, and the training images live in our own private storage. This step is only for a face you build yourself — presenters bought from the marketplace arrive already trained and are ready to use straight away.
Pick from a curated library of 15 voices across 5 voice styles, or clone a voice from a short clip you upload. The presenter keeps that voice everywhere.
Your presenter is now reusable. Talking-head videos are built from a still of your trained face, animated, then lip-synced to the voiceover — same person, every time.
Your trained presenter drives the parts of a video where a person is on screen delivering the message, plus opening shots built from a generated still of that face. Other generated scenes — b-roll, product shots, atmosphere — stay generic.
That's a limit of today's video models, which can't load a custom face model the way image models can. We could sprinkle your presenter's name into the b-roll prompts and call it a feature; it wouldn't change a single frame, so we don't. When video models support it, your existing presenter will be ready.
Videos featuring your trained presenter carry a small, dimmed label near the top of the frame. It's on by default.
Every script is generated under a standing rule that an AI presenter is never presented as a real customer, reviewer or testimonial-giver, and that customer quotes are never invented.
Requests to build a presenter described as a specific, identifiable real person are rejected before anything is generated.
If you clone a voice, you confirm you have the right to use that voice. Samples are used to create your presenter's voice, nothing else.
Full detail in our Terms and Privacy Policy.
You want a face fronting your content but you don't want it to be yours — and you don't want a different stock stranger in every video.
Consistency is what makes a face read as a brand. One trained presenter across every talking-head video beats six unrelated ones.
Give each client their own presenter and voice, kept apart, without booking a single shoot.
No. The face is generated, and the model is trained on that generated face — not on a real individual. We block attempts to build a presenter described as a specific, identifiable real person, and every script is held to a rule that an AI presenter is never passed off as a real customer, reviewer or testimonial-giver.
Yes. Videos that feature your trained presenter carry a small "AI generated" label on screen, on by default. It's sized to read as a compliance watermark rather than as part of the content.
Yes. You can clone a voice from a short clip you upload, or choose from the curated voice library instead. Either way the presenter uses the same voice in every video.
Talking-head segments and image-based opening shots. Other generated scenes stay generic, because today's video models can't load a custom face model — we'd rather say so than imply otherwise. See the note above.
You can build and keep multiple presenters, up to a per-account limit, and switch between them per video. Each one keeps its own face and voice.
Training is not retried automatically on purpose — a silent retry would double-bill a paid training run. You'll see the failure and can start it again yourself.
Tell us about your business and we'll get you set up.
Don't have an account yet? Start free — no credit card.
Rather see the characters first? Browse all 52 personas — you can buy one outright, no account needed.