AI Model Report

Reviews · SEPTEMBER 19, 2026

Salesforce's Koa Post-Trains Nemotron 3 Super for CRM Reasoning

Built via SFT and GRPO on synthetic scenarios spanning 14 industries, Koa matches or exceeds frontier models on Salesforce's own CRM benchmark with three times fewer errors — a claim no independent auditor has verified.

By Karl Strauchman · Senior model reviewer · September 19, 2026

Salesforce unveiled Koa at Dreamforce on September 15, 2026, its first CRM-specific reasoning model, post-trained with NVIDIA on the open-weight Nemotron-3-Super-120B and claiming three times fewer errors than leading frontier models on an internal benchmark. The framing is what matters: reasoning was the one capability Salesforce had until recently rented from the labs. Jayesh Govindarajan, the company's EVP of AI, described that arrangement to TechCrunch with two words. "Until now."

Read that beside the $63 billion fiscal-2030 revenue target Salesforce issued the following day, and Koa stops looking like a research artifact and starts looking like an infrastructure decision. Agentforce is the vehicle. Koa is what gets it off frontier-model rent.

The technical premise is straightforward. Salesforce and NVIDIA ran supervised fine-tuning followed by GRPO on synthetic scenarios spanning 14 industries, built from what the release describes as nearly three decades of CRM deployment patterns and using no customer data. The tooling stack was NeMo RL, NeMo Gym, and NeMo AutoModel; the agent runtime is Salesforce's declarative Agent Script. A September 14 arXiv paper cited by Unite.AI reports Koa improves on its Nemotron base across public tool-use and agentic-reasoning suites, with the clearest gains on multi-turn tool use. It beats a strong proprietary baseline. It still trails the strongest frontier models on general reasoning.

That's the honest read. The louder claim, the 3x-fewer-errors line, comes from a benchmark Salesforce designed and administered itself. TechTimes notes no independent auditor has verified it, which lands harder when you recall UC Berkeley's April 2026 finding that eight industry-standard agent benchmarks had been gamed to near-perfect scores without solving the underlying tasks, or Apple's GSM-Symbolic work showing up to 65% performance drops when problems are trivially reworded. Pilots inside Agentforce, at 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero, are the real proving ground. General availability slips to Winter 2026 in U.S. regions.

The structural argument, though, is the interesting one. Salesforce, holding 34.1% of the CRM market per Futurum, is betting that a model trained on the specific shape of lead qualification, opportunity updates, and case routing beats a generic model prompted toward the same tasks. This is the same premise our initial Dreamforce writeup traced last week, and it's the premise LemonLime builds on for small businesses: domain-specific sales and marketing work, prepared proactively, rather than a general model waiting for a prompt.

Govindarajan called Nemotron the first "sovereign American pre-trained model" with data provenance clean enough for the post-training run. The vendor benchmark won't survive independent scrutiny unchanged. The thesis underneath it might.

Sources

  • https://www.salesforce.com/news/press-releases/2026/09/15/koa-reasoning-model/
  • https://techcrunch.com/2026/09/15/salesforce-and-nvidias-new-reasoning-model-is-everything-the-ai-labs-should-fear/
  • https://www.cnbc.com/2026/09/16/salesforce-issues-revenue-target-of-63-billion-for-fiscal-2030.html
  • https://www.unite.ai/salesforce-debuts-koa-reasoning-model-for-agentforce-trained-on-nemotron/
  • https://www.techtimes.com/articles/327676/20260917/salesforce-launches-koa-crm-model-its-own-benchmark-agentforce-roi-gap-persists.htm