AI Model Report

Reviews · SEPTEMBER 20, 2026

Salesforce Ships Koa: A GRPO-Post-Trained Nemotron-120B That Routes CRM Reasoning Away From Claude and GPT

Unveiled at Dreamforce on September 15, Koa is Nemotron-3-Super-120B post-trained with Group Relative Policy Optimization on synthetic Agent Script workflows — and Salesforce's own paper puts it at 69.41 on Tau2Bench and 66.63 on BFCL.

By Karl Strauchman · Senior model reviewer · September 20, 2026

Salesforce introduced Koa at Dreamforce on September 15, 2026, its first in-house reasoning model for Agentforce, and the story worth telling isn't the demo. It's the routing decision. For workflows Salesforce controls, per-call trips to Anthropic and OpenAI are being replaced by a fine-tuned open-weight model the company runs itself.

Koa is Nemotron-3-Super-120B, NVIDIA's 120B open-weight base, post-trained by Salesforce Agentforce & AI Research using Group Relative Policy Optimization on synthetic Agent Script workflows generated across more than 14 industries. The technical report, dated September 14 and credited to Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, and Shubham Mehrotra, describes an unusually pragmatic training loop: NeMo Gym runs online rollouts, a frozen ~30B Nemotron-3-Nano v3 plays customer, tool emulator, and coverage judge, and the whole thing fits on five NVIDIA B200 nodes plus one helper. Serving is vLLM.

The numbers are legible. On Tau2Bench (50 airline, 114 retail, and 114 telecom tasks, four trials, pass^1 averaged, GPT-4.1 as user simulator), Koa scores 69.41, against 68.64 for the untuned Nemotron base and 54.48 for GPT-4.1, a 14.9-point margin. Claude Opus 4.8 posts 74.00 and GPT-5.5 posts 83.99. On BFCL, Koa hits 66.63% versus 64.73% for the base and 53.96% for GPT-4.1.

Then there's CRM Bench, which Salesforce designed. Koa lands at 0.86, behind GPT-5.5 (0.90) and Claude Opus 4.8 (0.87), ahead of GPT-4.1 (0.81) and the base (0.84). Post-training pushed function-call accuracy from 0.71 to 0.77. Salesforce's press-release headline, "3 times fewer errors" versus leading models on CRM actions, is graded on its own paper.

Jayesh Govindarajan, EVP of Salesforce AI, and Kari Ann Briski, NVIDIA's VP of Generative AI Software for Enterprise, are the named faces. The pilot roster (1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, Xero) is the intended signal that this ships. The adjacent AIforce layer exposes 60-plus MCP tools and 4,000-plus APIs, which is what a specialized reasoning model actually needs to be useful.

The skepticism is priced in. TechTimes surfaced UC Berkeley RDI research finding eight industry-standard agent benchmarks had been manipulated, and Apple's GSM-Symbolic work showed drops of up to 65% on minor problem variations. Constellation Research reads Koa as Salesforce edging toward the Palantir posture: own the reasoning layer, don't rent it. TechCrunch's framing, that this is what the labs should fear, is the tell. The frontier-API dependency was always a temporary arrangement, and application vendors with 27 years of workflow lore were always going to notice.