top of page

Tavus Announces Phoenix-4.5: A New State of the Art in Real-Time Human Rendering

Writer: Karan Bhatia
Karan Bhatia
1 hour ago
5 min read

Tavus, the human computing company, led by Hassaan Raza and Quinn Favret, has announced Phoenix-4.5, a new state-of-the-art in real-time human rendering.


Phoenix-4.5: More Natural Human Rendering.


Realism in digital humans goes beyond facial appearance. Natural movement, head motion, posture, shoulders, and reactions while listening is what make an interaction feel present.


Over three years, Phoenix has evolved from 3D conversational avatars to real-time facial and emotional animation. Phoenix-4.5 takes the next step by extending realistic movement across the head, shoulders, posture, and torso, creating more expressive and natural-looking PALs.


The model delivers this in real time with 134ms audio-to-video latency, which the company says is 25% faster than any other model on the market. 


Three Things Define Phoenix-4.5.


Phoenix-4.5 improves the three factors that make PALs feel more real: naturalness, identity, and responsiveness.


1. A major leap in naturalness

Phoenix-4.5 delivers richer facial animation, more expressive emotion, and natural micro-movements, while extending motion across the head, shoulders, posture, and torso. By generating the full frame rather than animating only the face, upper-body movement becomes part of the expression.


2. Stronger identity preservation

The model preserves the details that make a PAL look like its intended identity, even with more expressive animation. It also handles challenging features such as long hair, glasses, earrings, and headbands more reliably. In tests, 64% of faces that failed on Phoenix-4 successfully trained on Phoenix-4.5.


3. Faster real-time rendering

Phoenix-4.5 delivers audio-to-video rendering in 134ms, which the company says makes it the fastest real-time human rendering model on the market. It is also the first zero-shot Phoenix model, allowing new identities to be rendered from a single image or video almost instantly, with automatic refinement afterward.


From Face to PAL: Full-Frame Generation.


Phoenix-4.5 changes the underlying architecture by generating the entire frame in a single pass, face, head, neck, shoulders, torso, clothing, and surroundings, instead of treating the face as a separate layer.


Phoenix-4 previously generated a roughly 512×512-pixel facial region and composited it onto recorded body footage. This limited movement outside the face and could expose artifacts around the edges.


Phoenix-4.5 removes that boundary, allowing the entire upper body to move naturally with speech. The lips follow words, facial expressions follow meaning and intonation, while the head, shoulders, and torso respond to the rhythm of conversation.


The result is more than increased movement: coordinated movement across the whole PAL, making expression feel connected to how the character speaks and listens.


Human Behavior, Not Just Animation.


A PAL needs to remain expressive while listening, not just while speaking. Phoenix-4.5 makes these behaviors more natural, with richer expressions, more lifelike emotions, subtle micro-movements, and tighter lip sync.


Building on Phoenix-4’s active listening and continuous facial motion, the new model responds to both sides of the conversation, the PAL’s speech and the user’s input, to continuously determine how the PAL should move.


The result is more natural behavior across speaking, listening, reacting, and transitions, rather than switching between “talking” and “idle.”


Every Metric That Should Move, Moved.


Phoenix-4.5 shows measurable gains across key performance metrics. On the same internal test set, lip sync improved 16% over Phoenix-4, while motion vividness increased 1.3×. Identity fidelity remained broadly unchanged, while frame and video quality also improved.


Human preference testing shows an even clearer difference. In head-to-head Elo evaluations, zero-shot Phoenix-4.5 scored 1002, compared with 929 for fine-tuned Phoenix-4. Fine-tuned Phoenix-4.5 reached 1069 Elo.


The gains therefore extend beyond benchmarks: people can feel the difference in naturalness.


Creation Got a Lot Easier.


Phoenix-4.5 simplifies PAL creation in two major ways: new identities can be previewed almost immediately, and more people can create high-quality PALs.


Faster previews, less waiting

Phoenix-4.5 is the first Phoenix model to support zero-shot rendering, allowing a new identity to be previewed without upfront training. A preview is available in about one minute, while fine-tuning runs in the background. The finished fine-tuned model takes over after roughly two hours—half the time required by Phoenix-4.


The workflow shifts from wait, then see to see, then refine, with fine-tuning still improving detail and identity fidelity.


More people can become PALs.


Phoenix-4.5 also handles challenging real-world features such as hair, glasses, earrings, and other accessories more reliably because these regions are generated together rather than treated separately.


Of the identities that previously failed training on Phoenix-4, 64% trained successfully on Phoenix-4.5. Most remaining failures were linked to source-footage issues rather than the person’s appearance.


The result is a PAL creation process that is faster, less restrictive, and more inclusive.


Presence Changes What You Can Build.


Realism is not the end goal, the interaction is. In conversation, people constantly read signals beyond words: whether someone is listening, reacting, preparing to respond, and conveying emotion.


Phoenix-4.5 brings more of those cues to PALs, with natural movement, visible listening, and responsive expressions. This can make interactions in education, healthcare, coaching, sales, and support feel more like face-to-face conversations.


Early results from Tavus users show:

  • 2–3× higher engagement across sales and healthcare

  • 40% higher knowledge retention in learning and development

  • 50% faster ramp and 3× higher rep conversion in sales training


The goal is therefore not simply to make a Face look realistic, but to make a PAL feel human throughout the conversation.


Under the Hood: The Research.


Phoenix-4.5 combines two rebuilt models: an animator and a renderer.

The animator continuously processes both the PAL’s speech and the user’s voice, using a streaming diffusion transformer to generate compact motion signals conditioned on emotion, style, and identity. The model is distilled to just a few sampling steps, enabling real-time performance.


The renderer takes a reference image and those motion signals to generate the entire half-body frame jointly. Rather than compositing a moving face onto a static body, it generates coherent movement across the face, head, neck, and shoulders in a single pipeline.


Together, the models run in real time with under 130ms audio-to-video latency and around 35–40 frames per second end-to-end, including on slower machines.


Upgraded Existing Faces and New Faces.


All existing Tavus stock Faces have been upgraded to Phoenix-4.5, bringing improved naturalness, expression, and upper-body movement. The release also adds 50 new stock Faces, expanding the range of appearances, styles and personalities available out of the box.


The expanded library is designed to make it easier to find a production-ready Face for use cases ranging from tutors and coaches to sales, support and concierge PALs. Users can also create their own Face with Phoenix-4.5.


What Phoenix-4.5 Unlocks for Real-World Use Cases.


Phoenix-4.5 is designed for experiences where how an AI behaves on screen is part of the interaction, from sales and healthcare to interviews and training.


Sales Agents

AI sales agents need to maintain attention, respond naturally to objections, and stay engaging throughout a conversation. Tavus customers report 2× higher conversation engagement than previous chatbot tools and 6× more meetings booked.


Healthcare Agents

In patient intake, education, and follow-up, visible listening can make interactions feel more human. 70% of patients prefer face-to-face AI over text or audio for health conversations, while 97% of surveyed users in elder care preferred video over audio.


AI Interviewers

Realistic reactions can make screening and role-play more engaging and produce richer evaluation data. On Tavus, 75% of candidates return for a second session or more, while Colleva reduced hiring costs by 70% using AI video interviews.


Coaching and training

AI coaches need to respond to hesitation, maintain attention, and make practice feel like a real conversation. Tavus customers report 300% faster sales ramp-up and a 54% increase in confidence after practice sessions.

The same principle applies to concierge, support, and kiosk experiences: when the face is the interface, more natural rendering can make the entire interaction feel more human.


The Bar Is Face to Face.


Phoenix-4.5 is built around a simple standard: can AI feel natural in a face-to-face conversation?


The new model moves beyond a realistic-looking Face to a PAL that speaks, listens, reacts, and expresses itself as one continuous presence.


Phoenix-4.5 is now available for every PAL through the PAL Maker and API.


𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐌𝐞𝐧𝐥𝐨 𝐓𝐚𝐥𝐞𝐧𝐭 helps technology companies hire exceptional remote talent from India. Build your team with carefully curated engineers, product leaders, designers, GTM professionals, and more. Learn More At: https://www.menlotimes.com/menlo-talent

Menlo Times is a global media platform covering AI, Deeptech, Venture Capital, Fintech, Robotics, and Security through news, analysis, and insights from founders and operators.
  • Instagram
  • Facebook
  • X(Formerly Twitter)
  • LinkedIn
  • YouTube
© 2026 Menlo Times. All rights reserved.
bottom of page