top of page

Gimlet Labs Announces Series B

  • Writer: Karan Bhatia
    Karan Bhatia
  • 2 hours ago
  • 2 min read

Gimlet Labs, an applied research lab dedicated to building next-generation AI infrastructure, led by Zain Asgar, Michelle Nguyen, Omid Azizi, Natalie Serrino, James Bartlett, and the team, has announced $300M Series B raise, led by Andreessen Horowitz and joined by Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures, and XTX Markets.


The Need for Throughput and Speed.


Demand for AI tokens continues to accelerate, with monthly token generation up 6x in 12 months and projections pointing to another 20x increase by 2030.


Meeting this demand is driving nearly $1 trillion in annual AI infrastructure investment, with cumulative spending expected to reach $7 trillion by 2030. Power is emerging as the key bottleneck: AI data centers consumed roughly 18 GW in 2025, a figure expected to triple by 2030. Maximizing throughput per kW will be critical to scaling AI within physical resource constraints.


Beyond throughput, AI is creating growing demand for low-latency inference. Agentic workloads require repeated model calls, while models and context windows continue to grow. The challenge is no longer choosing between throughput and latency, but delivering fast tokens at high throughput.


Delivering Faster, More Efficient Inference with Gimlet Cloud.


Gimlet was founded on three beliefs: AI inference will become the dominant workload across software; current inference infrastructure is orders of magnitude less efficient than it could be; and meeting future demand will require infrastructure designed around inference from the ground up.


As inference overtakes training, existing training infrastructure has largely been repurposed to serve models. But inference has fundamentally different requirements, and those requirements vary significantly across workloads.


Meanwhile, a new generation of specialized AI accelerators is emerging, offering highly differentiated architectures and performance characteristics. Gimlet Cloud is built to harness these architectures for faster, more efficient inference.


At Gimlet Labs, the first multisilicon cloud is being built specifically for inference. The platform combines GPUs, near-memory compute, dataflow architectures, and CPUs, with software that distributes different inference tasks across the silicon best suited to each workload.


This approach can deliver 5–10x speedups at the same power footprint, or similar throughput gains at the same latency, critical as power becomes the key constraint on AI deployment.


Leveraging Heterogeneity with Granular Disaggregation.


Gimlet’s software decomposes models into components and dynamically schedules them across heterogeneous accelerators based on workload requirements and available capacity. This allows compute to be rebalanced across hardware, reducing stranded capacity and scaling individual inference tasks as demand changes.


The best-known approach is prefill/decode disaggregation, which runs compute-intensive prefill and memory-bandwidth-heavy decode on separate devices. Gimlet also supports techniques such as speculative-decode and attention-FFN disaggregation, offering different trade-offs between latency, throughput, and efficiency.


𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐌𝐞𝐧𝐥𝐨 𝐓𝐚𝐥𝐞𝐧𝐭 helps technology companies hire exceptional remote talent from India. Build your team with carefully curated engineers, product leaders, designers, GTM professionals, and more. Learn More At: https://www.menlotimes.com/menlo-talent

Menlo Times is a global media platform covering AI, Deeptech, Venture Capital, Fintech, Robotics, and Security through news, analysis, and insights from founders and operators.
  • Instagram
  • Facebook
  • X(Formerly Twitter)
  • LinkedIn
  • YouTube
© 2026 Menlo Times. All rights reserved.
bottom of page