How to create a fully realistic AI avatar with Higsfield
Most AI avatar workflows fall apart the moment you try to reuse the same character. The face changes between videos, the voice sounds different every time, and the result never quite feels like a real person. The good news: you can fix all of this with a structured workflow that keeps your character’s look and voice consistent across any number of videos—without filming anything or hiring anyone.
This guide walks through a six-step process inside Higsfield, an all-in-one AI platform for image, video, and audio generation. By the end, you’ll have a reusable avatar that looks and sounds like the same person in every scene.
Create a realistic character sheet (not a perfect model)
The foundation of a believable AI avatar is a strong reference image. Instead of generating a single portrait, start with a character sheet: one wide image that shows your character from multiple angles.
Inside Higsfield, go to the Image tab and create a new image. Set the model to Nano Banana Pro, the aspect ratio to 16:9, and the quality to 4K. This image will act as the master reference for everything you generate later, so higher quality here pays off in every future video.
Prompt the model to generate a three-panel character sheet: two full-body angles (front and back) plus a close-up portrait. Focus on realism over perfection. Include:
- Visible skin texture: pores, tiny marks, slight unevenness
- Imperfect details: stray hairs, slightly patchy eyebrows, a mole or beauty mark
- Distinct accessories: unique earrings, a specific hairstyle, or jewelry
- Flat, neutral lighting: no dramatic shadows or colored lights
These small imperfections are what make the character feel human. If you just ask for an “attractive woman in her mid-20s,” you’ll usually get an overly polished, almost plastic look that instantly reads as AI.
Save your character as a reusable element
Once you’re happy with the character sheet, save it so you can reuse this exact face and style in future generations.
In Higsfield, save the image as an Element and set the category to Character. From now on, you can reference this character in any prompt with the @ symbol, and Higsfield will pull directly from that sheet. This is how you keep your avatar’s appearance consistent across completely different videos and scenes.
Design props your avatar can interact with
Real people handle objects, show products, and interact with their environment. To make your avatar feel more grounded, generate props the same way you generated your character.
Back in the Image tab, switch the model to GPT Image 2, which is better for product-style images, and set the aspect ratio to 1:1. Prompt it to generate a simple, clear product—like a protein bar with bold text on the wrapper. You don’t need flashy design; clarity is more important than style for props.
Save this as an Element as well, but this time choose the Prop category. Now your avatar can reliably hold, show, or talk about this specific item in multiple videos.
Generate a casual talking-head video
With your character and prop ready, it’s time to test how well the avatar holds up in motion. Start with something simple and relatable, like a quick product review filmed as if it’s on a phone.
In the Video tab, use the Seance model (2.0 or 2.5 if available). Set the duration to around 15 seconds and use a vertical-friendly aspect ratio like 9:16. Keep the resolution at 1080p for now to save credits—you’ll upscale later.
Attach your Character and Prop elements under the prompt box, keep audio turned on, then write a detailed prompt that:
- Describes the setting (e.g., a small apartment kitchen, filmed on a selfie camera)
- Specifies the lighting (e.g., cool daylight from a window on one side, warmer kitchen light behind)
- Includes natural imperfections (slight grain, uneven lighting, handheld shakiness)
- Contains the exact spoken dialogue word-for-word
Letting the model invent its own dialogue usually leads to awkward, robotic speech. Writing the exact lines yourself gives you control over pacing and tone, while the model handles lip sync and performance.
Most importantly, lean into imperfections. Real handheld videos are rarely perfectly lit, framed, or stabilized. Slightly messy lighting, minor camera shake, and everyday environments all help your avatar blend in with real footage.
Build a location and shoot a travel-style vlog
Next, test whether your avatar still looks like the same person in a completely different context—like a travel vlog. For this, you need a reusable location reference.
Go back to the Image tab and switch the model to Higsfield Soul 2.0. Set the aspect ratio to 16:9 and prompt a detailed environment, such as a seaside terrace near a luxury resort. Include mixed materials and textures—old tiles, weathered concrete, wooden beams—to make the space feel lived-in.
Crucially, ask for the location to be empty of people. This image is a reference for the environment, not a finished photo, and you don’t want random background people that the model might regenerate differently every time.
Save this as a Location element. Location elements control the geography, materials, light direction, and overall atmosphere of a place, while your prompt will control framing and camera movement.
Then, in the Video tab, keep your model and duration the same but switch the aspect ratio to 16:9 for a standard horizontal vlog look. Attach both your Character and Location elements, and write a more structured prompt that breaks down the video into parts:
- How the camera moves (handheld, walking, slight shake)
- Where the character walks and what she reveals (e.g., turning a corner into the terrace)
- Her reactions and lines as she discovers the view
- How the environment should feel (lighting, time of day, weather)
For more control, you can break your prompt into blocks—one for the character, one for the camera, one for the environment, and one for audio. A useful block is a “positive locks” section: a list of things that must never change (face, hairstyle, earrings, general outfit style). This tells the model what needs to stay consistent even as it experiments with everything else.
If you’re interested in other ways to build vlog-style avatar content, you might also like this guide on using HeyGen to build realistic AI avatar videos.
Change outfits without changing the face
Real people don’t wear the same clothes in every video. To keep your avatar believable, you’ll want multiple outfits—without accidentally changing the face or overall look.
Instead of regenerating the character from scratch, go back to the Image tab with Nano Banana Pro and attach your existing Character element as a reference. Prompt a new character sheet that shows the same person in a new outfit, such as a cold-weather look with:
- A puffer jacket
- Rugged shoes and thick wool socks
- Nylon pants
Keep recognizable features like the hairstyle, earrings, and facial details intact. Save this new sheet as another Character element. Now you have the same person in multiple outfits you can call up for different scenarios—urban, travel, outdoor, professional, and more.
Create a cinematic multi-shot sequence
Once your avatar works in simple handheld videos, you can push things further with cinematic sequences. These are harder for AI because they combine multiple shots, cuts, and precise camera choices—but they also look the most impressive.
In the Video tab, keep Seance as your model and your duration similar, but change the aspect ratio to 21:9 for a widescreen, film-like look. Attach your cold-weather Character element and write a detailed prompt that describes a short sequence, for example:
- An opening drone shot over a cloudy mountain range at dawn
- A medium shot of your character stepping out of a tent with a mug
- A close-up as she takes a sip, with soft, natural morning light
- A final drone zoom-out showing the full campsite and landscape
Use optics to control your shots
To make the sequence feel like a real film shoot, include an “optics” block in your prompt that sets the field of view (FOV) for each shot. FOV is measured in degrees and determines how much of the world the camera sees:
- Wide shots: higher FOV (e.g., ~107°) to capture big landscapes and skies
- Medium shots: mid-range FOV (e.g., ~84°) for a natural human-eye feel
- Close-ups: low FOV (e.g., ~18°) to focus tightly on the face
This is especially important for close-ups, where your avatar’s face has to hold up under scrutiny. A lower FOV here helps keep the shot intimate and cinematic.
If you don’t pre-generate a location image for this scene, you can instead use a “location map” block in your video prompt describing the environment in detail—mountain type, vegetation, time of day, weather, and overall mood.
Add subtle, realistic audio
You can also add an audio block in the video prompt describing the background sound you want: soft ambient music, light wind, distant birds, or gentle atmospheric pads. This gives your sequence a cohesive mood without needing to edit audio separately afterward.
Design a custom voice that matches your character
Visual consistency is only half the story. If your avatar’s voice changes between videos—or sounds generic and robotic—it breaks the illusion. Higsfield lets you design a voice from scratch using text descriptions.
In the Audio tab, choose Text to Speech and set the model to Seed Audio 1.0. Leave the default settings as they are and don’t pick a preset voice. Instead, write a prompt that has two parts:
- A description of the person’s voice and delivery
- The exact lines you want them to say in this sample
For example, you might describe “a woman in her mid-20s with a light, warm timbre, a slight rasp, and a bit of vocal fry at the end of phrases, delivered in a bright and animated tone.” Then follow that with a short monologue your character might say in a vlog.
Seed Audio will generate a voice sample that matches your description. Because you’re specifying tone, texture, and energy level, the result is usually far more natural than picking a random stock voice and pasting in a script.
Turn the sample into a reusable voice preset
The audio file you just created can’t yet be attached directly to videos. To use it everywhere, you need to convert it into a reusable voice preset via voice cloning.
In the Audio tab, go to Voice Change. You’ll see two panels: one to pick a voice, and one to add your clip. First, click “Pick a voice” and upload the Seed Audio sample you just generated. Give it a name (for example, your character’s name), and Higsfield will create a custom voice preset from it. This takes about a minute and costs some credits, but you only need to do it once.
Now your custom voice appears alongside the stock presets and can be selected for any future voice change operations.
Replace the default video audio with your custom voice
With your voice preset ready, you can now upgrade the audio on your existing avatar videos.
Still in Voice Change, click “Add your clip” and upload one of your avatar videos, such as the kitchen product review. Select your custom voice preset, then generate the new audio. Higsfield will keep the timing and lip movements from the original video but replace the voice with your custom one.
Repeat the same process for your travel vlog and any other clips you’ve created. The result is that your avatar now has a consistent, natural-sounding voice across all videos—one that fits her age, personality, and style.
If you want to explore more approaches to lip-sync and voice consistency, you might also find this tutorial on creating 100% realistic AI lip-sync avatars from a single image helpful.
Upscale your videos to 4K without wasting credits
Generating everything in native 4K looks great but quickly becomes expensive. A more efficient workflow is to generate at 1080p, then upscale only the final videos to 4K.
In Higsfield, go to the Upscale section and upload your finished videos. Choose the Proteius model and select 4K as the target resolution. Rather than simply stretching the pixels, the upscaler adds real detail and sharpness, making your avatar videos look clean and ready to post on high-resolution platforms.
This approach lets you keep your day-to-day workflow affordable while still ending up with premium-quality output.
Build endless consistent avatar content
By combining a character sheet, reusable elements, structured prompts, a custom-designed voice, and 4K upscaling, you get a complete, repeatable system for AI avatar production:
- One face: locked in through character sheets and positive locks
- Multiple outfits: new sheets based on the same character reference
- Reusable locations: empty reference environments saved as Location elements
- Consistent voice: a custom preset built from a descriptive audio sample
- High-quality output: 1080p generation plus 4K upscaling
From here, you can create as many avatar videos as you like—product reviews, vlogs, cinematic sequences, tutorials—and your character will look and sound like the same person every time. All of it can be done entirely with AI, without cameras, studios, or actors.
Comments
No comments yet. Be the first to share your thoughts!