Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Nvidia Dynamo

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 2

  1. Nebius, NVIDIA and Uber on squeezing more value from every unit of compute

    At a panel hosted by The Information's AI Agenda Live, executives from Nebius, NVIDIA and Uber discussed how they are optimizing compute as agentic workflows threaten to multiply token consumption. Nebius CRO Marc Boroditsky said the company tailors optimizations in its Token Factory inference platform to each customer's goals—speed, reliability, quality or price—using techniques such as speculative decoding, where smaller models predict likely next tokens that a larger model verifies in bulk, tuned with the customer's own traffic. NVIDIA's Dion Harris pointed to a Dynamo feature that locates already-computed data in a cluster's KV cache so teams avoid recalculating it, while Uber infrastructure VP Mattie Toia said its token costs have stabilized despite growing agentic use, helped by caching pushed down to individual sub-agents and by giving engineers visibility into their own usage; Uber also built an internal "context graph" of roughly 24 million nodes that cut query times from 20 minutes to 30 seconds.

    The Information · 🔥 9
  2. Vast pitches tiered storage to ease AI agent memory pressure

    In an interview with theCUBE at CoreWeave's Fully Connected 2026 event, Vast Data co-founder and CTO Alon Horev said agent memory differs from ordinary inference: it includes both in-session context and long-term memory that lets an agent review past interactions. Vast's approach tiers storage — GPU memory first, then CPU memory on the same machine, then persistent media holding petabytes of KV cache — with Nvidia's Dynamo software orchestrating the process. Horev said a 500,000-token session can occupy one-tenth to one-twentieth of a GPU's memory, and offloading such sessions to storage avoids repeat recalculation; enterprises also need to record and retain everything their agents do, data that can feed fine-tuning or purpose-built models.

    SiliconANGLE AI · 🔥 5

Experience and discussion from the community

Share my Nvidia Dynamo experienceAsk about Nvidia Dynamo

Nobody has shared their experience with Nvidia Dynamo yet.