Soul App’s Real-Time Digital Human Stack: Beyond the Face to Fluent Social AI

Creating a digital human that truly feels "present" requires more than just realistic skin textures. It demands solving the "persistence problem"—keeping identity stable during long conversations—and mastering the subtle rhythms of human speech, from awkward pauses to overlapping interjections.

The Organisation for Economic Co-operation and Development (OECD) provides a wider frame through its AI Capability Indicators. Its social interaction scale treats interaction as an extended, embodied exchange between distinct participants. It identifies embodiment, social memory and identity as dimensions supporting communication, affective skills, social perception and social problem-solving.

Soul, an AI ecosystem company, shares a similar understanding of interaction. In a July 2026 interview, Tao Ming, Soul’s chief technology officer, described human conversation as a continuous process shaped by interruptions, speech breaks, filler words, tone and other paralinguistic signals. He also identified latency, visual quality, fluency, model size and cost as central constraints in real-time digital-human generation.

Years of product development on Soul App, its Chinese AI social platform, have helped Soul identify practical interaction problems and develop SoulX models across visual generation, speech processing, dialogue timing and expressive voice.

Generating a Coherent Digital Presence in Real Time

Real-time visual generation has to balance image quality, speed and stability. Soul developed SoulX-FlashHead and SoulX-LiveAct for different timescales of this task. FlashHead generates portrait video from streaming audio, where short input segments provide limited context and errors can accumulate over time.

Its Temporal Audio Context Cache retains information around each audio segment, while its training method reduces identity drift during longer sequences. The Lite version reaches 96 frames per second on a single NVIDIA RTX 4090.

LiveAct extends real-time human animation into much longer sequences. As generation continues, memory use grows, making it harder to carry stable visual information from one segment to the next. Neighbor Forcing improves consistency across adjacent frames, while ConvKV compresses earlier visual information into a fixed-length representation.

The model supports hour-scale animation and reaches 20 frames per second on two NVIDIA H100 or H200 GPUs. ConvKV preserves visual history during generation. Social memory operates at a broader level, maintaining continuity across past interactions and social context.

Following Speakers and Conversational Turns

In live conversation, a digital human needs to track who is speaking and when to respond. Similar voices, overlapping speech and rapid turn changes can leave a transcript without reliable speaker attribution. SoulX-Transcriber combines speech recognition with speaker diarization, the process of identifying who spoke and when. It links words and timestamps with the correct speaker, giving the system a structured account of how an exchange developed.

Managing conversational turn-taking is notoriously difficult.​ Unlike scripted dialogue, human speech is littered with hesitations and interruptions. SoulX-Duplug addresses this by distinguishing between a thoughtful pause and a finished utterance, allowing the AI to decide whether to jump in or stay silent.

Expanding the Expressive Range of Voice

Voice also shapes a digital character through tone, rhythm and style. SoulX-Singer extends this expressive layer into zero-shot singing synthesis. It generates singing from MIDI scores or melodic inputs, supports Mandarin Chinese, English and Cantonese, and provides control over pitch, rhythm and vocal style.

Facial realism gives a digital human a recognizable appearance. Social presence develops throughout the interaction: speech remains aligned with movement, the system follows who is speaking, turns occur at appropriate moments, and memory and identity preserve context across encounters.

SoulX addresses several of these engineering layers, while Soul App’s experience with social interaction provides the organizing context behind the portfolio. Soul is also bringing these capabilities into products and industry applications, including the B Soul intelligent device and scenario-specific solutions for industry partners.

By bridging the gap between high-fidelity graphics and nuanced social cognition, Soul is moving digital humans beyond mere "animations." With the SoulX stack, they are evolving into persistent social agents capable of genuine, long-term interaction—a critical step toward the next generation of embodied AI.