NVIDIA and CoreWeave Launch Vera Rubin for Agentic AI
At CoreWeave Fully Connected in San Francisco, CoreWeave announced the production availability of next-generation NVIDIA infrastructure, bringing advanced compute capabilities to high-demand enterprise environments.

Bringing Next-Generation Infrastructure to Production
Building on nearly a decade of close co-engineering, CoreWeave has integrated NVIDIA compute, networking, and software into a cloud platform purpose-built for artificial intelligence workloads. At the CoreWeave Fully Connected event running this week in San Francisco, the cloud provider announced the official availability of NVIDIA Vera Rubin NVL72 systems paired with Spectrum-X 102.4T Ethernet networking.
Cognition, the applied AI lab behind the Devin AI software engineer, serves as the first customer running production workloads on the newly deployed Vera Rubin architecture. Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, emphasized the longevity and flexibility of the platform, noting that older generations like NVIDIA V100 GPUs continue to support customer workloads alongside cutting-edge deployments.
Accelerating Agentic Workloads With Cognition
Cognition relies heavily on CoreWeave infrastructure to execute training, reinforcement learning, and production inference for Devin. Over a nine-month period, the applied AI lab scaled its operations to thousands of GPUs on CoreWeave Cloud to support demanding inference workloads.
Shortly after receiving its initial Vera Rubin NVL72 production racks earlier this month, Cognition benchmarked the hardware against a GB200 NVL72 baseline utilizing real-world software engineering tasks sampled from FrontierCode. Early tests demonstrated that the NVIDIA Vera Rubin NVL72 platform achieved up to a 4.8x increase in total token throughput for SWE-2 inference workloads over the previous baseline.
Silas Alberti of Cognition's founding team explained that agentic coding represents a uniquely complex workload characterized by long contexts, high concurrency, and massive token volumes where cost efficiency directly dictates deployment feasibility. According to Alberti, maintaining these workflows on a unified platform with collaborative engineering support outweighs any single hardware specification.
Expanding CoreWeave Cloud With NVIDIA Vera CPU
To address the dual infrastructure demands of agentic AI—serving large-scale low-latency compute and running thousands of isolated post-training environments concurrently—CoreWeave will also offer the NVIDIA Vera, marking the industry's first central processing unit built specifically for AI agents.
CoreWeave’s deployment packs 128 CPUs and 11,264 cores into a single rack, creating enough capacity to support over 11,000 concurrent environments configured at one core each. Utilizing BlueField-4 DPUs and Spectrum-X Ethernet switches, these hardware-isolated environments run securely alongside active training jobs while maintaining high performance and low latency.
In rigorous testing, CoreWeave recorded more than 3x faster agent sandbox startup times on the new CPUs. Furthermore, testing on Terminal-Bench revealed a 1.7x performance gain across all passing tasks, significantly accelerating the execution layers required for reinforcement learning, tool use, and model evaluation.
Closing the AI Loop With CoreWeave Forge
Historically, the iterative loop of model and agent improvement—where production behavior informs subsequent training runs and evaluations—has been fragmented across multiple vendor tools, leading to signal loss during handoffs. To solve this challenge, CoreWeave launched CoreWeave Forge, a unified environment integrating post-training expertise from OpenPipe, Weights & Biases, and the open-source marimo notebook project.
The Forge environment remains open across various models, frameworks, and cloud architectures. Key additions include CoreWeave ARIA for researching and iterating across the AI loop, CoreWeave Agent Lens for turning production traces into actionable failure-detection insights, and CoreWeave Sandboxes for secure, isolated code execution.
Enterprises can also leverage options such as serverless supervised fine-tuning and serverless RL to experiment with custom training recipes without necessarily requiring a dedicated training cluster.
Powering Next-Gen Inference and Enterprise Adoption
To maximize efficiency across enterprise deployments, CoreWeave incorporates NVIDIA Dynamo into its managed inference services and private preview rollouts. This open-source inference framework allows reinforcement learning rollouts to load new checkpoints directly into running deployments without requiring disruptive system redeployments.
Alongside specialized tools, enterprise builders such as Canva, Capital One, and MasterClass are actively utilizing the new cloud environment. Meanwhile, organizations can continue exploring foundational options using NVIDIA Nemotron open models to further customize their pipelines.
Sources
- NVIDIA BlogFrom Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI