@LinusEkenstam: I made a digital twin of myself from 10 seconds of video. In the clip: left is the real me, middle is a leading avatar …

X AI KOLs Following Products

Summary

Mirage Avatar X by Captions creates a realistic digital twin from just 10 seconds of video, preserving identity, expressions, and mannerisms without quality degradation.

I made a digital twin of myself from 10 seconds of video. In the clip: left is the real me, middle is a leading avatar model, right is Mirage Avatar X. Watch the eyes. The difference is not subtle. I have been testing AI avatar models since my first clone in 2023. Every one of them was impressive for about 30 seconds, then your brain caught up. Still eyes. One polite expression. A mouth doing all the work. Avatar X is the first model where that moment never came. Here is what makes it different: It is trained on you. Avatar X preserves your identity. Most avatar models can copy your appearance. Avatar X captures the subtle details that make you you. The way you move, the way you express yourself, and the way you naturally deliver speech. It looks like you. It moves like you. It sounds like you. It understands non-verbal performance Laughing, crying, yawning, sighing. These are the moments where most avatar models fall apart, trying to lip-sync through sounds that aren't words. Avatar X responds naturally, generating realistic facial expressions and micro-expressions instead of forcing every sound into speech. The expression goes beyond the lips Expressions are driven by the audio, through the whole face and body. Ask a question and it furrows its brows and shrugs on the tone. No other model does this to this degree. No quality degradation The first second and the last second look the same. Other models lose quality the longer the video runs. 10 seconds of input That is the entire requirement. Other models need 15 seconds, some even 1 to five minutes. Three years ago my AI clone was a party trick. This one can carry my face, my expressions and my delivery without me in the room. The bar for AI avatars just moved. Avatar X is live today. → Try it here: https://captions.ai/get/ai-twin?utm_source=x&utm_medium=influencer&utm_campaign=avatarx&utm_ad=linusekenstam…
Original Article
View Cached Full Text

Cached at: 07/28/26, 04:36 PM

I made a digital twin of myself from 10 seconds of video.

In the clip: left is the real me, middle is a leading avatar model, right is Mirage Avatar X.

Watch the eyes. The difference is not subtle.

I have been testing AI avatar models since my first clone in 2023. Every one of them was impressive for about 30 seconds, then your brain caught up. Still eyes. One polite expression. A mouth doing all the work.

Avatar X is the first model where that moment never came.

Here is what makes it different:

It is trained on you. Avatar X preserves your identity.

Most avatar models can copy your appearance. Avatar X captures the subtle details that make you you. The way you move, the way you express yourself, and the way you naturally deliver speech.

It looks like you. It moves like you. It sounds like you.

It understands non-verbal performance

Laughing, crying, yawning, sighing. These are the moments where most avatar models fall apart, trying to lip-sync through sounds that aren’t words.

Avatar X responds naturally, generating realistic facial expressions and micro-expressions instead of forcing every sound into speech.

The expression goes beyond the lips

Expressions are driven by the audio, through the whole face and body. Ask a question and it furrows its brows and shrugs on the tone. No other model does this to this degree.

No quality degradation

The first second and the last second look the same. Other models lose quality the longer the video runs.

10 seconds of input

That is the entire requirement. Other models need 15 seconds, some even 1 to five minutes.

Three years ago my AI clone was a party trick. This one can carry my face, my expressions and my delivery without me in the room.

The bar for AI avatars just moved. Avatar X is live today.

→ Try it here: https://captions.ai/get/ai-twin?utm_source=x&utm_medium=influencer&utm_campaign=avatarx&utm_ad=linusekenstam…


Ten seconds from now, there can be two of you.

Source: https://captions.ai/get/ai-twin?utm_source=x&utm_medium=influencer&utm_campaign=avatarx&utm_ad=linusekenstam

Your digital twin that actually looks and sounds like you.

Now you can create more content, reach more people, and show up consistently with your AI twin. Powered by Mirage Avatar X, our most advanced model yet.

Create your twinCreate your twin

Micro-expressions

Laughs like you laugh

Positive:Laughs, then resumes the video

Negative:Keeps mouthing “words” through the laugh

Non-verbal emotion

Reads the silence

Positive:Natural pause with living micro-motion

Negative:Dead-eyed idle loop

Quality under pressure

No degradation mid-take

Positive:Same fidelity at 0:59 as 0:01

Negative:Mouth shapes degrade; lips stop matching audio

Your recording

Mirage Avatar X only needs 10 seconds of footage to make your twin.

Your twin

Your AI twin shares your full identity, voice and mannerisms.

10 seconds of footage

Captions

15 seconds of footage

HeyGen*

1-5 minutes of footage

Synthesia*

Avatar X

-10-second video reference -Holds likeness across the full video -Consistent quality from first frame to last -Full emotional range -Natural blinks, glances, and subtle motion -Learns from your reference footage

Previous model

-Single photo reference -Drifts from the source photo -Degrades as the video progresses -Limited, flatter delivery -Micro-expressions not supported -Inferred, often generic

How to make a digital twin in Captions

Record yourself

Record yourself

Record a short video, reading the provided script. You can do this in theCaptions apporonline.

Add a script

Add a script

Enter the script you want your twin to deliver in a new video.

Create the video

Create the video

Generate new videos in minutes, and export directly to top platforms.

Frequently asked questions

What do I actually need to record?One continuous 10-second clip, filmed directly in Captions. Avatar X extracts your expression range, voice character, and motion from that single take. You can optionally add a few reference photos to sharpen quality.

What should I say in my recording?We give you a short script designed to capture your natural expression range. Just read it the way you’d talk to a friend.

Do I need to add extra photos?No, Captions creates your twin from the clip alone. You can add other photos if you want to, but it’s not necessary.

Is my twin only mine?Yes. Twins are consent-first. You have to agree that you have consent and the right to make a twin before you create it.

How is this different from other avatar models?Most platforms need minutes of footage, and their avatars break on anything non-verbal like laughing, pausing, or reacting. Mirage Avatar X builds a fully expressive twin from 10 seconds, and holds frame-accurate lip-sync.

Does my twin sound like me too?Yes. Our Mirage Audio model builds a voice clone directly from your 10-second video. It captures your inflection, tone, and delivery too.

Do I need to re-record to change how my twin looks?No. Your 10-second clip is a one-time setup. After that, you can create a Look from a prompt or photo with no camera required.

What video formats are supported?Horizontal and vertical, so the same twin works for YouTube, ads, Reels, TikTok, and Shorts. You can make videos up to 3 minutes long.

What languages are supported?Your twin can speak 30+ languages thanks to AI dubbing and translation. You can also add captions or subtitles in 100+ languages.

How are Captions twins different from HeyGen or Synthesia twins?Most platforms need minutes of footage, and their avatars break on anything non-verbal like laughing or pausing. Avatar X builds a fully expressive twin from 10 seconds of footage and maintains top-notch quality across your entire video.

*From each platform’s published requirements.

Similar Articles

Google Makes It Easy to Deepfake Yourself

Wired

Google has introduced a new 'avatar' feature in its Flow tool, allowing users to create digital clones of themselves and insert them into AI-generated videos using the Omni Flash model.

Avatar V: Scaling Video-Reference Avatar Video Generation

Hugging Face Daily Papers

Avatar V is a production-scale framework for generating behaviorally recognizable avatar videos conditioned on full video references, introducing sparse reference attention and motion representation streams to achieve state-of-the-art identity preservation and lip synchronization.