AI image models reinvent a face every time you generate. The moment you want the same person or character to appear across many images and clips, you need an avatar: a reusable identity you attach to a generation so the result looks consistent. Higgsfield gives you two ways to make one, and choosing the right one is the whole skill. This article explains both — Soul training and Reference Elements — when to use each, and how avatars carry through the rest of the platform.

Why avatars exist

Consistency. Without an avatar, “a portrait of my character” produces a different face on every call. An avatar pins one identity so the character in image one is the character in image twenty, and in the video and the 3D model after that. The two mechanisms below solve this in different ways, with a real trade-off between fidelity and flexibility.

Path 1: Soul training — a faithful identity from photos

A Soul is a trained identity model of one person. You give Higgsfield a name and 5 to 20 reference photos, it trains for about ten minutes, and you get a soul_id that reproduces that exact face.

  • How you make one. Through the Soul Characters tool (show_characters with the train action): supply the name and 5–20 clear photos of the same person. Training runs asynchronously; you check status until it is ready.
  • How you use one. Generate with a Soul model and pass the soul_id: generate_image with model: "soul_2" (Soul 2.0) or soul_cinematic (Soul Cinema). Every generation then carries that identity.
  • The constraints that matter. A trained Soul works only with the Soul models (Soul 2.0 and Soul Cinema), and there is one identity per generation — you cannot put two Souls in the same shot.

Soul is the right call when the priority is a faithful likeness of one real person: your digital twin, a recurring on-screen host, identity-consistent portraits, fashion, or cinematic stills.

Path 2: Reference Elements — instant, flexible references

A Reference Element is a reusable reference saved from an image — a character, but also an environment or a prop. There is no training: you create it from one or more images and it is ready immediately.

  • How you make one. Through the Elements tool (show_reference_elements with the create action): pass one or more images and it returns an element id, synchronously.
  • How you use one. You reference the element by name inside your prompt for generate_image or generate_video, and Higgsfield injects the saved image behind the scenes. Because you can reference several elements in one prompt, you can put two characters — or a character and a specific location and a prop — in the same shot.
  • Which models. Elements work with a broad set: Nano Banana Pro and Nano Banana 2, GPT Image 2, Seedream 4.5 and 5.0 lite, Cinema Studio Image 2.5, Cinema Studio Video, Seedance 2.0, and Kling 3.0. They do not work with the Soul models.

Elements are the right call when you want speed, more than one subject in a shot, a non-person subject (a place or an object), or when you simply want to use one of the non-Soul models.

Which one to use

The two systems are complementary, not competing. Pick by the shot:

QuestionSoul trainingReference Element
Faithful likeness of one real person?Yes, this is its jobApproximate
Two or more characters in one shot?No — one identity per generationYes — reference several
A place or a prop, not a person?No — people onlyYes
How long to create?About 10 minutes (trains)Instant
Reference images needed5–20 of one personOne or more
Which models it works withSoul 2.0 and Soul Cinema onlyNano Banana, GPT Image 2, Seedream, Cinema Studio, Seedance, Kling

A useful rule of thumb: train a Soul for a person you will feature repeatedly and alone; save an Element for everything else — multi-subject scenes, locations, props, or a quick one-off reference on your model of choice.

Inventing a character with no photos

If you do not have photos because the character is invented, the Soul family has a text-driven option: Soul Cast generates a consistent cinematic character from a description alone. It is the “make me a persona” path, versus Soul training’s “clone this real person” path.

Avatars for ads

There is a third, ad-specific avatar concept in Marketing Studio. Its avatar library holds presenters — preset or custom — that you pair with a product to make talking-head and UGC-style ads. It is a separate system from Soul and Elements, aimed at “a person holding and pitching my product,” and it is covered by the marketing side of Higgsfield rather than the general identity tools here.

The flow in practice

Training and using a Soul, conversationally:

Train a Soul character named “Ada” from these eight photos of me.

Higgsfield starts training (about ten minutes) and returns a soul_id once ready. Then:

Generate a studio portrait of Ada in soft window light.

generate_image
  model: "soul_2"
  soul_id: "<the trained soul_id>"
  prompt: "studio portrait in soft window light, shallow depth of field"

Creating and using an Element, for a two-character scene a Soul could not do:

Save this character as a reusable element, then put her and a friend in a Paris cafe, using Nano Banana Pro.

Higgsfield saves the element, and the generation references it (and a second one) in a single Nano Banana Pro image.

Cost and time

  • Soul training charges a one-time training fee and takes roughly ten minutes; after that, each generation costs whatever the Soul model costs.
  • Elements are instant to create; you pay only when you generate with them, at the underlying image or video model’s price.

As always, preview a specific generation’s cost with the estimate before committing, which spends nothing.

How avatars carry through the pipeline

An avatar is most valuable because it survives every downstream step. A consistent face from a Soul or an Element flows into:

Because identity is set at the image stage, the habit is the same as everywhere in Higgsfield: get the character right in a still first, then carry it into video and 3D.

Tips

  • Feed a Soul varied photos. Five to twenty clear shots across angles and lighting train a more faithful identity than a handful of near-identical selfies.
  • One person per Soul. For a scene with two people, use Elements — a Soul is a single identity per generation.
  • Match the model to the method. A trained Soul needs a Soul model (soul_2 or soul_cinematic); an Element needs a non-Soul model (Nano Banana, Seedream, Kling, Cinema Studio, Seedance).
  • Use Elements for places and props, not just people — a saved location or object keeps scenes consistent too.
  • Invent with Soul Cast when there is no real person to photograph.

Recap

Higgsfield gives you one goal — a consistent identity — through two mechanisms. Soul training builds a faithful model of one real person from 5–20 photos in about ten minutes and works with the Soul models, one identity per shot. Reference Elements save a character, place, or prop instantly, let you combine several in one prompt, and work across the non-Soul models. Train a Soul for a person you will feature repeatedly and alone; save an Element for multi-subject scenes, non-people, and quick references — and either way, set the identity in a still, then carry it into video and 3D.