@LinusEkenstam: I made a digital twin of myself from 10 seconds of video. In the clip: left is the real me, middle is a leading avatar …
Summary
Mirage Avatar X by Captions creates a realistic digital twin from just 10 seconds of video, preserving identity, expressions, and mannerisms without quality degradation.
View Cached Full Text
Cached at: 07/28/26, 04:36 PM
I made a digital twin of myself from 10 seconds of video.
In the clip: left is the real me, middle is a leading avatar model, right is Mirage Avatar X.
Watch the eyes. The difference is not subtle.
I have been testing AI avatar models since my first clone in 2023. Every one of them was impressive for about 30 seconds, then your brain caught up. Still eyes. One polite expression. A mouth doing all the work.
Avatar X is the first model where that moment never came.
Here is what makes it different:
It is trained on you. Avatar X preserves your identity.
Most avatar models can copy your appearance. Avatar X captures the subtle details that make you you. The way you move, the way you express yourself, and the way you naturally deliver speech.
It looks like you. It moves like you. It sounds like you.
It understands non-verbal performance
Laughing, crying, yawning, sighing. These are the moments where most avatar models fall apart, trying to lip-sync through sounds that aren’t words.
Avatar X responds naturally, generating realistic facial expressions and micro-expressions instead of forcing every sound into speech.
The expression goes beyond the lips
Expressions are driven by the audio, through the whole face and body. Ask a question and it furrows its brows and shrugs on the tone. No other model does this to this degree.
No quality degradation
The first second and the last second look the same. Other models lose quality the longer the video runs.
10 seconds of input
That is the entire requirement. Other models need 15 seconds, some even 1 to five minutes.
Three years ago my AI clone was a party trick. This one can carry my face, my expressions and my delivery without me in the room.
The bar for AI avatars just moved. Avatar X is live today.
→ Try it here: https://captions.ai/get/ai-twin?utm_source=x&utm_medium=influencer&utm_campaign=avatarx&utm_ad=linusekenstam…
Ten seconds from now, there can be two of you.
Your digital twin that actually looks and sounds like you.
Now you can create more content, reach more people, and show up consistently with your AI twin. Powered by Mirage Avatar X, our most advanced model yet.
Create your twinCreate your twin
Micro-expressions
Laughs like you laugh
Positive:Laughs, then resumes the video
Negative:Keeps mouthing “words” through the laugh
Non-verbal emotion
Reads the silence
Positive:Natural pause with living micro-motion
Negative:Dead-eyed idle loop
Quality under pressure
No degradation mid-take
Positive:Same fidelity at 0:59 as 0:01
Negative:Mouth shapes degrade; lips stop matching audio
Your recording
Mirage Avatar X only needs 10 seconds of footage to make your twin.
Your twin
Your AI twin shares your full identity, voice and mannerisms.
10 seconds of footage
Captions
15 seconds of footage
HeyGen*
1-5 minutes of footage
Synthesia*
Avatar X
-10-second video reference -Holds likeness across the full video -Consistent quality from first frame to last -Full emotional range -Natural blinks, glances, and subtle motion -Learns from your reference footage
Previous model
-Single photo reference -Drifts from the source photo -Degrades as the video progresses -Limited, flatter delivery -Micro-expressions not supported -Inferred, often generic
How to make a digital twin in Captions
![]()
Record yourself
Record a short video, reading the provided script. You can do this in theCaptions apporonline.

Add a script
Enter the script you want your twin to deliver in a new video.

Create the video
Generate new videos in minutes, and export directly to top platforms.

Frequently asked questions
What do I actually need to record?One continuous 10-second clip, filmed directly in Captions. Avatar X extracts your expression range, voice character, and motion from that single take. You can optionally add a few reference photos to sharpen quality.
What should I say in my recording?We give you a short script designed to capture your natural expression range. Just read it the way you’d talk to a friend.
Do I need to add extra photos?No, Captions creates your twin from the clip alone. You can add other photos if you want to, but it’s not necessary.
Is my twin only mine?Yes. Twins are consent-first. You have to agree that you have consent and the right to make a twin before you create it.
How is this different from other avatar models?Most platforms need minutes of footage, and their avatars break on anything non-verbal like laughing, pausing, or reacting. Mirage Avatar X builds a fully expressive twin from 10 seconds, and holds frame-accurate lip-sync.
Does my twin sound like me too?Yes. Our Mirage Audio model builds a voice clone directly from your 10-second video. It captures your inflection, tone, and delivery too.
Do I need to re-record to change how my twin looks?No. Your 10-second clip is a one-time setup. After that, you can create a Look from a prompt or photo with no camera required.
What video formats are supported?Horizontal and vertical, so the same twin works for YouTube, ads, Reels, TikTok, and Shorts. You can make videos up to 3 minutes long.
What languages are supported?Your twin can speak 30+ languages thanks to AI dubbing and translation. You can also add captions or subtitles in 100+ languages.
How are Captions twins different from HeyGen or Synthesia twins?Most platforms need minutes of footage, and their avatars break on anything non-verbal like laughing or pausing. Avatar X builds a fully expressive twin from 10 seconds of footage and maintains top-notch quality across your entire video.
*From each platform’s published requirements.
Similar Articles
@trymirage: Introducing Mirage Avatar X, the new standard for AI avatar models. Avatar X delivers industry-leading identity preserv…
Mirage Avatar X is a new AI avatar model that claims industry-leading identity preservation, expressiveness, and support for vertical and horizontal video.
I Cloned Myself With Gemini’s AI Avatar Tool. The Result Was Unnervingly Me
The author tests Google's Gemini AI avatar tool, which creates a digital clone from a selfie video to insert into AI-generated videos, finding the result impressively realistic yet unsettling.
Google Makes It Easy to Deepfake Yourself
Google has introduced a new 'avatar' feature in its Flow tool, allowing users to create digital clones of themselves and insert them into AI-generated videos using the Omni Flash model.
@DanKornas: Want to generate avatar videos without sending your face and voice data to a cloud service? Duix.Avatar is a local AI a…
Duix.Avatar is a local AI avatar toolkit that generates lip-synced avatar videos from video, voice, and script inputs without sending data to the cloud. It supports eight languages and is available under a community license on GitHub.
Avatar V: Scaling Video-Reference Avatar Video Generation
Avatar V is a production-scale framework for generating behaviorally recognizable avatar videos conditioned on full video references, introducing sparse reference attention and motion representation streams to achieve state-of-the-art identity preservation and lip synchronization.