Home / Blog / Why Capital One built its multi-agent AI platform around…
Tech News

Why Capital One built its multi-agent AI platform around open-weight models

I'll pull the VentureBeat source and the post-writing rules first so the 600–900 words stay factual and match the house format.The request is a standalone…

By Dillip Chowdary • Aug 14, 2026 • Source: VentureBeat

Why Capital One built its multi-agent AI platform around open-weight models

What happened

I'll pull the VentureBeat source and the post-writing rules first so the 600–900 words stay factual and match the house format.The request is a standalone news body, not an engineering deep-dive. Next I’ll pull the full VentureBeat piece and the local news-post pattern so the paragraphs stay factual.VentureBeat is behind a bot wall, so I’ll pull the piece from the local news ledger and any cached copies instead of inventing details.The prompt matches the post-writer prose slot, so I’ll check that contract and a recent live body, then write the 600–900 words from the verified VentureBeat facts only.The user asked for the post body itself, so I’ll write the 600–900 words from the verified VentureBeat facts only—no invented numbers, no markdown.At VB Transform 2026, Kel Vanee, MVP of machine learning engineering at Capital One, spoke with Sam Witteveen, Senior Technology Contributor at VentureBeat, about how the bank built a scalable multi-agent AI architecture around deeply customized open-weight models rather than an off-the-shelf foundation model. Vanee said Capital One is not just using AI, it is building AI. He traced that stance to earlier work on data transformation and cloud adoption, which he said let the bank move when the current AI wave arrived. Those investments produced three decisions: a centralized, enterprise-wide AI platform with built-in governance, open models customized on proprietary data, and a multi-agent orchestration harness Capital One constructed itself.

Capital One fine-tunes open-weight models on data Vanee called a huge advantage that nobody else has and that general frontier models cannot provide. Real-time data is treated as required, because live customer and associate interactions need fresh context. Customization is not siloed. Vanee said that as the bank customizes open-source models for one use case, it sees benefits across the whole portfolio, because the training makes the model an expert at Capital One use cases, policy, and nomenclature, and that training produces a general lift. Chat Concierge, the bank's customer-facing auto-shopping assistant, runs on a version of Meta's open-weight Llama model that has been customized with that same proprietary data.

The technical detail

Why Capital One built its multi-agent AI platform around open-weight models
Illustration · Pexels

The production example is a bank-fraud customer-service workflow that handles millions of calls a year, with interactions that run from roughly four minutes to as long as sixty minutes. Sending those calls through a single large language model was not enough. Capital One's multi-agentic workflow, MACAW, routes work through specialized agents with governance and guardrails built in. An understanding agent looks at what the customer is saying and tries to recover intention. A reasoning agent is given specific instructions and writes a summary. A validation agent fact-checks that summary. An explaining agent turns it into a formatted document that is shared with human agents. The workflow supports several hundred customer-service agents who specialize in complex fraud calls, and the post-call summaries replace reconstructing long, back-and-forth conversations by hand.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Why it matters for builders

Chat Concierge uses the same split: one agent converses with the customer, one builds an action plan from business rules, one evaluates accuracy, and one explains and validates the result. That is the builder-relevant claim. A single model that interpreted, planned, checked, and wrote the artifact failed on the fraud-call workload; putting a validation step between specialized agents is how the bank made the system usable for several hundred specialists. The same stack also runs off the customer path. Capital One built an autonomous agentic optimization solution that tunes backend hosting infrastructure so researchers can search for latency without executing every experiment themselves.

Vanee said the research system walks a search space the researcher designs, handles the mechanics of setting up and running each experiment, and then puts a summarization of the results in front of the human. He noted that new LLM optimizations arrive constantly and are not all complementary: combining two good optimizations can cause a performance regression. That is the market contrast with renting one hosted frontier model and accepting vendor updates as they come. Capital One is optimizing its own hosting and its own customized open-weight stack so it can hit the quality, cost, and latency thresholds it set for itself. Vanee said those ingredients were absolutely critical to differentiating in this space.

Market and competitive context

What he said to watch next is routing, not a bigger single model. He described routing abstraction layers that a platform validates over multiple models for cost and accuracy, and argued that routing across a broader set of available models can produce better accuracy than any individual model because different models excel in different areas. The other shift he flagged is proactive, event-driven AI that steps in when it detects conditions that warrant action, rather than waiting for a human prompt. He tied that specifically to larger-scale monitoring and to fighting fraud, and he said deploying those agents would demand rigorous testing and monitoring.

What to watch next

The open questions sit next to those claims. MACAW's validation agent is the answer to a single model's failure on long fraud calls, but the conversation does not say how often the validator disagrees with the reasoner or what happens when the understanding agent misreads intention on a sixty-minute call. Chat Concierge is said to use a customized version of Meta's Llama; the article does not name the Llama generation. Combining optimizations can regress latency, which is why the agentic search loop exists, but that loop still depends on a researcher-designed search space. Proactive agents that act without being asked raise the same governance problem the centralized platform was built to contain. Those are the constraints to track as Capital One expands the platform across more use cases for its millions of customers.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →