How Nuance Labs is Building the Foundation Model for Human Conversation and Expression


Nuance Labs, building a human foundation model with emotional intelligence, led by Fangchang Ma, Edward Zhang, and Karren Yang, has raised $50 million in Series A funding led by returning investor Lightspeed Venture Partners, with participation from returning investors Accel, South Park Commons, and new investors NVIDIA and Define Ventures. The company will open a research preview of its model to the public later this year.
Nuance Labs is developing AI that perceives and responds to human communication in real time, combining words and tone with facial expression, gaze, gesture, and timing. Transcript-based AI captures only part of that interaction, leaving important signals about intent and meaning unseen.
By interpreting these multimodal cues and responding expressively, Nuance Labs aims to make AI interactions more natural and intuitive. When AI can better understand and reflect human communication, users can express themselves more freely, allowing more of their intent and meaning to come through.
Why Face-to-Face AI Has Fallen Short.
Voice AI has gained traction because it is more useful than earlier interfaces. Face-to-face AI has struggled because systems rely on disconnected components for transcription, reasoning, speech, and animation, creating latency and breaking the flow of conversation.
Nuance Labs uses a specialized full-duplex model that perceives and responds simultaneously, allowing AI to process incoming signals and generate responses in real time while a conversation is still underway.
The Nuance Labs Approach: One Full-Duplex Model for Human Conversation.
Nuance Labs’ model:
Interprets words, tone, gaze, gesture, and timing.
Responds in real time through facial and vocal expression.
Uses a single full-duplex model to perceive and respond simultaneously.
Signals understanding through verbal and non-verbal cues while the conversation is still underway.
“The potential for AI and avatars to enhance our lives will never be achieved while they stare at you blankly or keep interrupting you. They've never felt natural because we've been contorting ourselves and how we communicate to the machine rather than the machine adapting to us,” said Fangchang Ma, Co-founder and CEO of Nuance Labs. “The most productive collaboration comes from being able to express yourself freely, in words, tone, gesture, and expression, the way you would with a friend or close colleague, with all the nuance in the back-and-forth that turns talking into understanding. That's what we're building at Nuance Labs: AI that understands the many ways we express ourselves and responds the way a person does, in the moment.”
The Founding Team: From Apple Research to a Single-Model Approach.
Nuance Labs’ founding team spent years at Apple researching how machines perceive and reconstruct people, revealing the difficulty of accurately capturing human presence and interaction. Even telepresence remained challenging, particularly when accounting for the verbal and non-verbal cues of live conversation.
Nuance Labs was founded to address this challenge with a single model that natively understands and generates human conversation and expression from audio and video, rather than stitching together separate systems for transcription, language, speech, and animation.
How the Nuance Labs Foundation Model Works.
Nuance Labs’ foundation model learns from how people communicate and operates in full-duplex fashion, simultaneously streaming audiovisual perception and generating audiovisual responses. Real-time interaction is therefore built into the architecture rather than added as a separate layer.
The model interprets motion, expression, and meaning as they unfold and can respond while a person is still speaking. The approach is designed to make AI conversations more natural and effective, particularly in settings where human expression and communication directly influence outcomes.
Potential applications include sales and customer service, coaching, professional training, education, and other environments where AI works alongside people.
“There's a massive opportunity to fix the way people interact with these transformative AI tools. Nuance Labs is building that interface, one that can see and respond to a person the way people do,” said Nnamdi Iregbulem, Partner at Lightspeed Venture Partners. “The founding team has a rare mix of technical talent and operating excellence, and they've turned that into a single model that follows you in real time. We can see this becoming a foundational layer for AI products everywhere, and this is the team to build it."
The funding will accelerate model development, expand the research team, and support Nuance Labs’ first public research preview. The preview will allow users to interact face-to-face with AI that interprets audiovisual cues and responds in real time, bringing more natural human interaction to AI.
𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐌𝐞𝐧𝐥𝐨 𝐓𝐚𝐥𝐞𝐧𝐭 helps technology companies hire exceptional remote talent from India. Build your team with carefully curated engineers, product leaders, designers, GTM professionals, and more. Learn More At: https://www.menlotimes.com/menlo-talent


