CoreWeave Forge connects production data to AI agent improvements
A new platform aims to close the gap between operations teams monitoring healthy services and AI engineers trying to improve model quality. CoreWeave announced Forge at its Fully Connected 2026 conference to help teams trace failures back to fixes and prove improvements before deployment.

Operational dashboards often mask deteriorating agent performance. The solution requires converting production failures into owned evaluation tasks that teams can act on systematically.
Once an AI agent reaches production and the service remains stable, the real challenge emerges: Is answer quality improving? How can teams connect production traces, curated datasets, and evaluations to spot problems, implement solutions, and validate new versions before users encounter them?
Getting an initial deployment live marks progress. But sustaining it introduces harder problems: maintaining accuracy, keeping customers satisfied, and managing costs. After several cycles, teams often spend more effort on tool maintenance and coordination than on advancing the application.
Bridging organizational silos
When AI engineers hand off a trained model to application or site reliability engineering (SRE) teams, a structural problem emerges. The researchers know what quality looks like; the production team manages the running system. Yet they often examine different signals.
Application traces flow into dashboards designed for uptime monitoring. The on-call engineer sees a healthy service, while the research team gains little visibility into answer quality or customer sentiment. Evaluation suites become stale as user behavior shifts, but no one takes responsibility for updating them. Researchers, AI engineers, and business stakeholders return to separate dashboards for the next release.
The on-call engineer sees a healthy service, while the research team learns little about answer quality or user sentiment.
The starting point is identifying who can convert a production failure into an evaluation case and who maintains that case over time. If answering these questions requires moving context between teams, that's where coordination breaks down.
The five-stage improvement cycle
The AI loop follows a repeating pattern: deploy a model or agent, monitor its behavior, transform signals into data, modify the system, test the result, and start again. Each stage must preserve enough information for the next team to proceed.
Run
Select the model and agent harness—the surrounding code and infrastructure—for the task. Gather the signals needed to investigate performance after the system goes live.
Observe
Look past uptime metrics. Collect traces, performance data, tool usage, and user feedback to examine agent decisions and pinpoint failure points.
Curate
Convert production instances into training datasets and refreshed evaluation benchmarks. Keep human judgment in the loop and maintain records showing where each example originated.
Improve
Align the change with the identified problem. This might mean adjusting the harness, switching models, or tuning behavior through reinforcement learning (RL), supervised fine-tuning, or model distillation. Specify the expected outcome in quality, speed, or expense.
Evaluate
Test the candidate against consistent benchmarks before, during, and after rollout.
A release should show what improved, rather than rely on a few promising answers.
Keeping records with the work
These handoff challenges shaped the design of CoreWeave Forge, unveiled last week at the Fully Connected 2026 conference. The platform integrates running, observing, curating, improving, and evaluating within a single development workspace. MasterClass and Canva are beginning to adopt Forge to complete their AI loops.
A meaningful test for an integrated environment is whether teams can trace a release back to its model version, evaluation suite, dataset, and production examples. CoreWeave Registry manages models, agents, and datasets; Weights & Biases Models tracks experiments, hyperparameter sweeps, analysis, and automated workflows. These records enable teams to understand what changed and whether results improved.
For monitoring, CoreWeave Agent Lens traces steps, decisions, and tool invocations, offering conversation views and technical breakdowns. Live traffic scoring incorporates human review and curation. According to the launch announcement, Agent Lens improves failure detection by 20% and reduces fix costs by half; teams should validate these claims against their specific workloads.
CoreWeave Notebooks supplies managed Python environments for shared evaluations and custom analysis. CoreWeave ARIA examines runs, suggests experiments, and recommends code changes stored in GitHub. Its connection to Weights & Biases Models enables automated research based on detected patterns. Evaluation evidence stays alongside proposed changes so teams can assess outcomes.
CoreWeave Sandboxes delivers isolated central processing unit (CPU) or graphics processing unit (GPU) environments for agents, tool calls, RL, and evaluations. CoreWeave ARIA and CoreWeave Sandboxes are generally available; Agent Lens, Notebooks, and Model Distillation are in preview and newly part of CoreWeave Forge.
Selecting the right improvement approach
Post-Training leverages production signals to enhance model quality, latency, and cost without requiring a training cluster. Serverless SFT (supervised fine-tuning) and Serverless RL (reinforcement learning) enable teams to experiment with their own training methods.
For tasks already validated in production, Model Distillation trains a smaller, open-weights model on outputs from the larger model. It then compares the candidate directly against the current version. This comparison provides evidence for deciding whether to shift traffic. A smaller model delivers value only if it meets the task's requirements.
Inference as part of the system
Every stage depends on inference, but different workloads demand different control levels. CoreWeave Inference provides Serverless Inference for accessing open-weights models and Dedicated Inference for workloads on isolated hardware.
Serverless Inference puts infrastructure management on CoreWeave. Dedicated Inference gives teams control over model weights, deployment settings, and GPU allocation. Selection depends on workload needs; teams can switch between options as requirements change.
Cline uses open-source models through Serverless Inference for its open-source, open-choice coding agent, which the launch post describes as trusted by more than 11 million developers. Grammarly uses Dedicated Inference with explicit GPU selection and fully managed operations.
RL introduces additional complexity. The policy generates rollouts—samples of its behavior—training produces updated weights, and those weights return to serving for the next cycle. Slow checkpoint loading becomes a bottleneck.
The new RL rollouts feature in Dedicated Inference uses the NVIDIA Dynamo foundation to hot-load updated checkpoints with minimal downtime. CoreWeave partnered with NVIDIA and you.com's engineering team to post-train NVIDIA Nemotron 3.5 Lightning using RL Rollouts in NeMo gym with you.com's web search application programming interface (API). The launch post reports a 15-fold improvement in model reload latency compared to baseline.
Integrating with existing infrastructure
An improvement loop extends into observability platforms, data warehouses, and security systems already in place. The CoreWeave Partner Network encompasses infrastructure, data services, independent software vendors (ISVs), and models and inference, with solutions tested on CoreWeave under production conditions.
This includes tools that agents invoke. Exa, Parallel Web Systems, and You.com provide live web search through a single integration. Search gives agents access to information outside their training data; teams still need to monitor and evaluate tool calls and resulting answers.
CoreWeave Forge was built for teams running workloads on CoreWeave Cloud, other cloud providers, and on-premises data centers. Improvement data remains portable, with open interfaces between stages, so curated datasets and evaluation suites stay useful regardless of deployment location.
Getting started
Begin with a single production failure your team can clearly understand. Track it from trace through curated example, proposed solution, evaluation, and deployment. Document where context gets lost and where accountability blurs. This provides a concrete foundation for building the loop.
CoreWeave Forge offers a 30-day Pro free trial with credits across product lines for teams wanting to test the workflow. For those seeking to validate production workloads with direct access to CoreWeave experts and infrastructure, CoreWeave ARENA is available. The objective is ensuring each deployment informs what gets built next, backed by evidence that the next version outperforms the previous one.