Industry

Cognition's Devin and CoreWeave Forge Build Infrastructure for Continuous AI Agent Learning

AI agent systems are shifting toward continuous loops of inference, feedback and training rather than discrete releases. Cognition AI and CoreWeave are building infrastructure to support always-on learning at scale across distributed hardware.

3 min read
Always-on AI agents turn infrastructure into a continuous learning loop

The architecture supporting AI agents is undergoing fundamental change to accommodate systems that cycle perpetually through inference, feedback collection and model training. Cognition AI Inc.'s Devin product now spans the entire software development workflow—from initial planning and code generation through code review and production incident response. Managing this expanded scope demands infrastructure designed to sustain continuous learning across large distributed systems, according to Silas Alberti, head of research and founding team member at Cognition.

Our runs are always on. While we still ship releases, package them up a little bit, I think the reality is we're always training. We're always trying to find the next data and the next reward signals to improve our models.

Silas Alberti, Cognition AI

Alberti and Chen Goldberg, executive vice president of product and engineering at CoreWeave Inc., discussed these infrastructure challenges during the Fully Connected event, speaking with theCUBE Research's Dave Vellante and John Furrier in an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio.

Always-on training raises the reliability bar

The convergence of training and inference has become especially pronounced in reinforcement learning scenarios, where model outputs generate the feedback signals needed for improvement. Cognition operates distributed training infrastructure spanning multiple data centers across different countries and continents, making consistent uptime across thousands of graphics processing units a critical requirement, Alberti explained.

If you run a big training run, you're not just looking at the GPUs, but you also want uptime. If just one replica goes down, the whole training run goes down. I think a very important metric is getting to this 99.99% reliability.

Silas Alberti, Cognition AI

Hardware advances are also steering Cognition's research priorities. Early access to Nvidia Corp.'s Vera Rubin platform gives the company's researchers the opportunity to examine the system and refine model designs while targeting improved price-performance ratios.

With each generation, price performance just goes up, so we can do more with the same amount of compute. Being early and actually being able to study the kernels and the dynamics of this new platform allows us to prioritize our research investments.

Silas Alberti, Cognition AI

How CoreWeave Forge connects the AI agent infrastructure loop

CoreWeave unveiled Forge during the event as a platform designed to integrate inference, observation, data curation, model refinement and evaluation into a unified system. The offering provides Agent Lens for monitoring agent behavior, alongside capabilities for model distillation and reinforcement learning. A feature called RL Rollouts enables updated model checkpoints to be loaded into production systems without requiring a full system redeployment.

CoreWeave Forge is … tailored for those AI loops, those learning systems of how you build agents and connect all the different steps. You don't have to do everything at once. You can always start with serverless inference, find your model and start.

Chen Goldberg, CoreWeave

For Cognition, the continuous learning loop creates an opportunity to leverage real-world operational experience to enhance future iterations of Devin. The objective is to expand the agent's effectiveness across extended software development projects, according to Alberti.

https://www.youtube.com/embed/u0elJ0DQj6w?feature=oembed

Agents are not just planning the code, writing the code [and] reviewing the code, but also becoming real production maintainers. If there's a production issue, Devin can jump on it, respond to it. And by the time you wake up, there's already a [pull request] and you can merge it.

Silas Alberti, Cognition AI

Source: SiliconANGLE · Reporting supplemented by The Silicon Ledger staff.