AI & DevelopmentCloud & DevOpsDeveloper Tools

CoreWeave Forge: The AI Dev Platform Built Into Your GPU Cloud

CoreWeave Forge unified AI development platform with ARIA and Agent Lens

CoreWeave spent the last two years becoming the dominant GPU cloud. Now it wants to own everything above the silicon too. On September 30, the company launched Forge — a unified AI development platform that wraps training, inference, evaluation, experiment tracking, and agent observability into a single loop. The GPU landlord just handed you a lease on the whole building.

The Problem Forge Actually Solves

If you’ve shipped an AI product, you know the pain: your production traces don’t feed your next training run. Your finished experiments don’t inform your next evaluation. Every handoff between tools loses signal and costs time. The average AI team cobbles together four to six tools to cover what Forge claims to handle natively.

CoreWeave calls its architecture a five-stage loop: Run → Observe → Curate → Improve → Evaluate — and then repeat. The idea is that improvement compounds across model versions instead of resetting with each new training run. Nick Patience of Futurum Group framed it concisely: “Closing the loop from production back into training in a single environment, while staying open across models, frameworks, and clouds, is a harder engineering problem than either half alone.”

CoreWeave didn’t build this from scratch. It acquired Weights & Biases (MLOps), OpenPipe (fine-tuning), and integrated the open-source Marimo notebook project. Forge is the product that wraps them together with new CoreWeave-native services.

ARIA: Your AI Research Agent Is Now GA

The headline feature is ARIA (AI Research and Iteration Agent), which shipped to general availability at launch. ARIA lives inside Weights & Biases and behaves like a research collaborator who has read every experiment you’ve ever run.

The interaction model is conversational. You ask something like “Does adding dropout after attention layers reduce overfitting?” and ARIA formulates a hypothesis, writes the configuration files, launches the experiment through W&B Launch, evaluates the results against your baselines, and proposes follow-up runs. It doesn’t generate text summaries — it builds shareable W&B workspaces with appropriate visualizations: heat maps for parameter sweeps, parallel coordinates for hyperparameter interactions, bar charts for discrete comparisons.

ARIA has access to your full project context from the start: training code, experiment logs, loss curves, artifacts, and checkpoints across multiple projects. It’s available on mobile through the W&B app. Persistent memory and MCP server support are on the roadmap.

This isn’t a chatbot wrapper bolted onto your toolchain. ARIA operates inside your existing infrastructure and executes — not just advises.

Agent Lens: The Feature That Actually Matters for 2026

ARIA gets the headlines, but Agent Lens — currently in public preview — is arguably the most timely addition in the Forge stack. As teams deploy autonomous agents at scale, “what did my agent actually do and why did it fail?” has no clean answer today. Agent Lens traces every step, decision, and tool call end-to-end.

The numbers CoreWeave cites: 20% improvement in critical failure detection, and issue resolution at one-tenth the cost compared to using a general-purpose frontier LLM for debugging. Those figures matter when an agent is taking real actions in production.

Agent Lens also converts production signals into curated datasets — failures become training data automatically. That’s the production-to-training loop made concrete.

What’s Available Today

Before planning the migration, here’s what’s actually production-ready:

  • Generally Available: ARIA, Sandboxes, Weights & Biases Models, Serverless SFT, Serverless RL
  • Public Preview: Agent Lens, RL Rollouts
  • New (limited access): Model Distillation, CoreWeave Notebooks

Pricing: free for individuals, $60/month for teams under 50, custom enterprise. There’s a 30-day Pro trial with real compute credits, so you can stress-test it before committing.

The Lock-In You Should Know About

When your GPU provider also ships your observability platform, training pipeline, and agent sandbox, switching becomes a compound cost — not a single migration. CoreWeave says the right things: Registry stores assets in open formats, Forge works across models and clouds. But they’re also planning NVIDIA Vera CPU integration that promises 3x faster agent sandbox startup versus x86. That’s a performance moat, not a contractual one.

Forkast’s analysis captured it well: “switching costs are structural, not contractual.” You won’t be trapped. You’ll just have a very good reason to stay.

The Bottom Line

Forge is the most coherent attempt yet to close the gap between “model in production” and “better model next week.” For teams already on CoreWeave GPU infrastructure, the free tier is a no-brainer to try. For teams on other clouds, the question is whether the loop integration justifies the gravitational pull toward a single vendor.

ARIA is genuinely novel. Agent Lens is genuinely underbuilt everywhere else. The lock-in is real but subtle. If you’re building AI products at scale, Forge is worth an afternoon of evaluation — just go in with eyes open.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *