How Mundo is Building the Data Layer for Perceptual Intelligence
- Karan Bhatia

- 28 minutes ago
- 2 min read

Mundo, building the data layer for perceptual intelligence, led by Jason Liao, Naijide Anwaer, Garreth Lee, and Kenneth Wu, has raised a $20 million Series A led by GreatPoint Ventures, with participation from Y Combinator, Next Frontier Ventures, and E12 Ventures. Together with its previously unannounced $4 million seed round, Mundo has now raised $24 million.
The Missing Half of Intelligence.
AI has made extraordinary progress in recent years. Models can write software, solve complex mathematical problems, answer difficult questions, and reason through increasingly sophisticated tasks.
But intelligence requires more than reasoning. It also requires perception.
Humans constantly interpret signals beyond words: facial expressions, tone, gestures, timing, body language, background context, and changes in the environment. A conversation can shift meaning based on a pause or expression that never appears in a transcript.
AI systems are becoming increasingly capable of reasoning over structured information, but much of the real world remains unstructured. Conversations overlap, background noise disrupts speech, cameras capture incomplete views, and people communicate through subtle signals that are difficult to represent as data.
Understanding these signals represents a fundamentally different challenge.
Perceptual intelligence is the ability for AI to understand the richness of real-world sensory experiences, from speech and video to gestures, environments, and human interaction, and use that understanding to interact naturally. It could become a defining capability of the next generation of AI systems.
From Internet-Scale Data to Real-World Experiences.
The first generation of foundation models learned from internet-scale data. The next generation will increasingly learn from richer forms of human experience, including audio, video, and emerging modalities. These require not only more data, but new types of data and new ways to measure progress.
Existing benchmarks often fail to capture the capabilities needed for real-world interaction. New datasets and evaluations are therefore becoming essential for advancing multimodal AI, from natural speech-to-speech interaction and fine-grained video understanding to emerging modalities without established learning methods.
The core principle is that AI progress will depend on more than better models. Better evaluations and better data must evolve together in a continuous feedback loop, shaping how the next generation of perceptual intelligence is built and measured.
𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐌𝐞𝐧𝐥𝐨 𝐓𝐚𝐥𝐞𝐧𝐭 helps technology companies hire exceptional remote talent from India. Build your team with carefully curated engineers, product leaders, designers, GTM professionals, and more. Learn More At: https://www.menlotimes.com/menlo-talent


