Reviews · SEPTEMBER 18, 2026
Reviewed: Salesforce's Koa Is a Post-Trained Nemotron Reasoning Model Built to Displace Frontier Routing Inside Agentforce
Post-trained on NVIDIA Nemotron 3 Super with 27 years of synthetic CRM scenarios, Koa claims three times fewer errors than leading models on Salesforce's own CRM Bench — on a benchmark no independent party has yet verified.
Salesforce introduced Koa at Dreamforce on September 15, 2026, and the framing is unambiguous: this is the CRM giant's first in-house reasoning model, post-trained on NVIDIA's Nemotron 3 Super, and it exists to stop routing Agentforce workloads to Claude and ChatGPT at frontier prices. The company claims Koa "matches or exceeds leading model performance on CRM actions with three times fewer errors" on its own CRM Bench suite. No independent party has verified that number.
The architectural thesis is coherent even before the benchmarks are litigated. Salesforce says it post-trained Nemotron on the equivalent of 27 years of synthetic CRM scenarios spanning more than 14 industries, using NVIDIA's NeMo RL, NeMo Gym, and NeMo AutoModel toolchain and Group Relative Policy Optimization for reinforcement learning. The output is a hybrid Mixture-of-Experts model with a one-million-token context window, activating only a subset of parameters per request. Trending Topics reports the accompanying paper claims 11 percent more precision on action selection, 2.1x greater reliability, and 15 percent better long-context retention against general-purpose models.
Jayesh Govindarajan, Salesforce's EVP of AI, told TechCrunch the origin story plainly: "there was no sovereign American pre-trained model that was available… state of the art, and… had clear data provenance." Read that as procurement logic dressed as strategy. Agentforce runs at $2 per conversation, and every reasoning step routed to a frontier lab is margin ceded to a supplier Salesforce doesn't control.
The counter-evidence is the benchmark itself. CRM Bench is vendor-administered, and TechTimes flags the UC Berkeley work on benchmark manipulation and Apple's GSM-Symbolic findings as reasons to treat single-source scoring as marketing until reproduced. The pilot list, 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero, will produce more legible signal than any leaderboard chart.
Salesforce is also hedging. Claudeforce keeps Anthropic's model available as an interface for the up to 1,000 clients CNBC reports have signed up for the Claude beta, and Missionforce Nemotron is due in October ahead of Koa's winter 2026 U.S. general availability. The $63 billion fiscal-2030 revenue target announced this week is the number all of this has to service.
The strategic read: Salesforce is doing what every application-layer incumbent with distribution eventually does once frontier inference becomes a variable cost line rather than a capability moat. Bloomberg-terminal logic, applied to CRM. The companies that already covered this pattern from the model side, Anthropic's Fable 5.1 benchmark leap and what GPT-6 Astra actually shipped, are now the suppliers Koa is designed to route around.
Sources
- https://www.salesforce.com/news/press-releases/2026/09/15/koa-reasoning-model/
- https://techcrunch.com/2026/09/15/salesforce-and-nvidias-new-reasoning-model-is-everything-the-ai-labs-should-fear/
- https://www.cnbc.com/2026/09/16/salesforce-issues-revenue-target-of-63-billion-for-fiscal-2030.html
- https://www.techtimes.com/articles/327676/20260917/salesforce-launches-koa-crm-model-its-own-benchmark-agentforce-roi-gap-persists.htm
- https://www.trendingtopics.eu/salesforce-builds-its-first-ai-model-on-nvidias-open-weight-model-nemotron/