Skip to main content

How a deepfake video call runs

Everything on this page is true whatever platform you run the call on. The six channel pages cover only what differs: how you get into the call, and what goes wrong there.

A deepfake video call is operator run. Someone on your team joins a real call wearing a real-time face and voice swap, and works the scenario live. It is not an automated call that runs while you do something else, and the operator's attention is the scarce resource in the whole exercise. Budget for it.

How you get into the call is your decision

Callstrike puts an operator into a video call wearing someone else's face. It does not decide whose call, or how the operator got there. That is a judgement you make per exercise, and it changes the exercise far more than anything in the wizard:

  • Host the meeting yourself and invite the targets. You control the time, who is in the room, and when the reveal happens. Nobody is surprised into a call they did not agree to attend. This is the usual shape for awareness training, and the easiest to run cleanly.
  • Join a meeting you were invited to. The operator has a legitimate invite, usually because the exercise was arranged with someone inside the team. The pretext is that the persona is simply attending a meeting they were expected at.
  • Join a meeting you were not invited to. The operator has a code or link obtained the way a real attacker would. This is red teaming, it is the most realistic, and it is the one that needs authorisation in writing before it happens, naming who may be targeted and what is out of scope.

The three differ in who consented to what, not in how the product works. An awareness session where everyone knows a deepfake is coming and a red team engagement where nobody does are both legitimate, and both are run from this same screen. Decide which you are doing, write it down, and make sure the person who authorised the exercise agrees with your reading.

⚠ Whichever shape you pick, the target is a real person who will be recorded reacting to a deception. In several jurisdictions the disclosure is not optional. See AI Act compliance.

Before you start

  • The feature has to be enabled for your workspace. A trial workspace can build a campaign and save it as a draft, but pressing Launch shows an upgrade prompt instead of starting a session.
  • One live session at a time, per workspace. Deepfake video calls and voice impersonation calls draw on the same slot. If one is running, the second attempt is refused, and only the admin who started a session can use it.
  • Two operators cannot share a session. If someone else in your workspace has one open, you wait for them to finish.
  • Three channels need an account and a phone. WhatsApp, Telegram and Signal are joined from an account you already hold, paired from a mobile device in the room. The other channels need neither.

1. Build the campaign

Open Deepfake Video Call in the console sidebar and choose Create New Campaign. The wizard is three steps.

The Deepfake Video Call campaign list in the Callstrike console, showing six campaigns with the identity being impersonated, the use case, the target department, the delivery platform and each campaign's status
The Deepfake Video Call campaign list, where an operator starts.

Delivery Channel

Pick the platform. Zoom, Microsoft Teams, Google Meet, Webex and GoTo Meetings offer a place to paste the meeting's joining details; Slack, WhatsApp, Telegram and Signal do not, because there is nothing to paste.

The new deepfake campaign wizard on its Delivery Channel step, with Zoom, Slack, Microsoft Teams, Telegram, Signal, WhatsApp, Google Meet, Webex and GoTo Meetings offered as selectable cards
Step 1 of the wizard: the platform the call will run on.

Those details are one free text note, and they are optional. It is not a link field and it is not validated: whatever you type is shown back to you as a floating note during the call, so you are not hunting through a calendar invite with a live meeting on screen. Paste the meeting ID and passcode, or the code, or nothing.

Changing the channel clears the note. A warning bar says so once you have entered something. It is deliberate: a Zoom meeting ID is not a Meet code, and carrying one across would give you a note that is confidently wrong.

Scenario Setup

Four groups of fields, and between them they are the exercise:

  • Exercise configuration. The use case (tabletop exercise, red team test, awareness training) and the audience. The use case is what the report is framed as afterwards.
  • Impersonated identity. The name and job title the operator will wear. Someone the target knows and has reason to defer to.
  • Attack configuration. The target department, the objective (a wire transfer, credentials, data, access, a vendor payment redirect), and the pressure tactics in play: urgency, authority, familiarity, confidentiality, and so on.
  • Generated scenario. A drafted name and description built from everything above. Regenerate it, or write over it. It is a starting point, not an output.
The new deepfake campaign wizard on its Scenario Setup step, showing the use case and target audience fields, the name and job title of the person being impersonated, and a panel of operator instructions
Step 2 of the wizard: the exercise, the audience, and the identity.

Scenario Summary

Read it back, then Launch. The browser asks permission to send notifications, because the next part takes long enough that you will not sit and watch it.

2. The wait

Launching provisions a dedicated session for your workspace. It usually takes about four minutes. When capacity is short the platform tries other regions, and the wait can reach about ten.

The Callstrike console provisioning a deepfake session, showing the wait it promises and a progress bar partway through step two of the launch
Launching a session. The wait is usually about four minutes and can reach ten.

Those numbers are measured, not estimated, and the spread is real rather than a safety margin: about half of launches are quick, and about a quarter take the long path. Plan the exercise around the long one. If you have told a target to expect a call at ten past, launch at the hour.

Do not close the tab. Leaving the page tears the session down. If it has already produced a recording the session is ended; otherwise it is cancelled outright. There is no pause.

3. Configure the session

Three numbered steps, in order. You can go back to any of them from the same screen.

Step 1: Audio Configuration

Choose the voice the target will hear. This is the persona's voice, not yours: your speech is converted on the way out.

Test it before you go any further. The voice is the half of the impersonation people notice first, and a voice that does not fit the face and title you have chosen is worth catching now rather than in the first ten seconds of a call.

Step one of a deepfake session, Audio Configuration, with a reference voice attached and the control that tests the conversion, before the conversion is started
Step 1: the reference voice, ready to test.

Step 2: Video Setup

Upload the face the target will see, and start the swap. Activation takes roughly half a minute to a minute, and the preview shows you the result before anyone else sees it.

Skipping the face swap does not fall back to your own camera. It leaves no camera at all. The swap is what produces the video the meeting application picks up, so with it off the application finds nothing to send and the target sees a blank tile. Audio is separate and is unaffected. If you continue past this step without an active swap, the console warns you, and the warning means what it says.

Step two of a deepfake session, showing the operator's camera feed beside the same feed with the persona's face swapped in, and the swap marked active
Step 2: the swapped result is visible to the operator before anyone else sees it. The person shown is a model, not a Callstrike employee.Footage filmed for Callstrike; persona face supplied by Callstrike (Filmed for Callstrike with permission. Not stock footage)

Step 3: Ready to Launch

The last check before the meeting application opens. After this you are looking at the channel, and the channel pages take over.

4. The call

The meeting application fills the screen, with a control rail beside it. The application is running on the session, not on your machine: you drive it with your mouse and keyboard, and it has no address bar and no tabs.

The first time you use a channel, a short guide opens over it pointing at the one control that gets you into the meeting. It can be reopened at any time from the session footer.

Everything above is your side of the call. This is the other side, which is the side that decides whether the exercise works: the persona is simply another participant, on camera, under whatever name was typed at join.

A Google Meet call as the other participants see it, with the persona on camera under the name it joined with, beside a participant whose camera is off
What the room sees: the persona, on camera, under the name typed at join.Footage filmed for Callstrike; persona face supplied by Callstrike (Filmed for Callstrike with permission. Not stock footage)

The rail

The rail is what separates a simulation from an impersonation. It is how you stop being the persona, on purpose, in front of the target:

  • My real face. Drop the swap, either instantly or as a four second transition. The transition is the one to use when the point is that the target watches it happen.
  • My real voice. The same, for the voice. The two are independent, so you can keep the face and drop the voice, or the reverse.
  • Split screen and side by side. Show both faces at once, labelled, rather than replacing one with the other.
  • AI disclosure. Put a notice on the outgoing video.
The reveal rail beside a live deepfake session, listing the real face and real voice controls, the split screen and side by side controls, and the AI disclosure setting
The reveal rail: every way an operator can drop the persona, and the disclosure notice.

Turning the disclosure on changes what everyone else receives, not what you see. It is burnt into the outgoing video, so it survives into the recording and cannot be cropped out by the person watching:

The same Google Meet call with the AI disclosure turned on, showing a labelled box around the persona's face and a caption reading AI manipulated video
The same call with the disclosure on. This is what the notice looks like to the people receiving it.Footage filmed for Callstrike; persona face supplied by Callstrike (Filmed for Callstrike with permission. Not stock footage)

Decide before the call which of these you are going to use and when. A reveal that is fumbled live is worse for the debrief than one that is scripted, and in several jurisdictions the disclosure is not optional. See AI Act compliance.

If the call is being recorded, a REC marker sits on the video, and it is the same marker that is burnt into the recording.

5. Afterwards

End the session from the footer. Then, on the campaign, record what actually happened: the outcome, a vulnerability level, and a written observation. There is help available to draft the summary from the scenario and the outcome, which is worth using while it is fresh.

The campaign feedback form after a deepfake call, with the outcome, the target vulnerability rating and a written observation being recorded
After the call: the outcome, the vulnerability rating, and what the operator observed.

Recordings appear on the campaign's detail page, with a thumbnail and duration each. That page is what you take into the debrief.

What goes wrong, on any channel

The launch button will not start a session. Either the workspace is on a trial, or a deepfake or voice impersonation session is already running for it.

The wait ran past ten minutes. The session is still hunting for capacity. This is the tail of a real distribution rather than a fault, but if you are on the clock with a target, it is a reason to reschedule rather than to keep waiting.

The target sees a black tile. The face swap is not running. Go back to Video Setup and start it. This is the single most common cause, and the reason the warning at step 2 exists.

The voice is yours. The voice reveal is on, or the voice was never started at step 1.

The session vanished mid call. The tab was closed or navigated away from. There is no recovery: the session ends, and anything already recorded is kept.

Pick your channel

Microsoft Teams, Webex and GoTo Meetings are supported and work the same way as Zoom; they do not yet have a page of their own. The deepfake video call reference covers all nine.