Article
→ Open-weight models are now production-grade for agentic workloads. Automation Anywhere's benchmark across 150+ enterprise agent scenarios found NVIDIA Nemotron 3 Super outperforms "almost all large closed-source models" evaluated — the first major automation platform vendor to publicly validate this performance threshold. (Automation Anywhere, 2025)
→ Data sovereignty, not model capability, is the primary driver of on-prem AI adoption. Regulated industries — financial services, healthcare, insurance, government — are converging on fully self-hosted AI architectures because data residency mandates leave no alternative, regardless of cloud model performance.
→ A reference architecture is rapidly commoditising. Accenture, EY, and Deloitte are independently converging on the same on-premises enterprise AI automation platform stack: NVIDIA AI Enterprise software + Dell-class infrastructure + open reasoning models — signalling that the infrastructure layer is becoming a commodity, and competitive differentiation will shift to agent orchestration and governance. (Accenture newsroom, 2025; EY newsroom, 2025)
→ Governance is the critical gap, not technology. Only 1 in 5 organisations has a mature model for overseeing autonomous agents — yet agentic AI usage is "poised to rise sharply" in the next two years. The bottleneck for on-premises enterprise AI automation is organisational readiness, not compute. (Deloitte, State of AI in the Enterprise, 2026)
Why This Matters Now
For four years, the on-premises enterprise AI automation platform was a compromise architecture — organisations accepted it for compliance reasons while conceding that proprietary cloud models simply performed better. That trade-off has collapsed.
In the first half of 2025, three independent forces converged simultaneously: an open-weight model that matches or exceeds closed-source benchmarks on real enterprise agent workloads; a hardware and software reference stack endorsed by the world's largest professional services firms; and a wave of data sovereignty regulation that is transforming on-prem from a risk-management preference into a legal requirement.
The scale of the shift is measurable. Worker access to AI rose 50% in 2025, and the number of companies with 40% or more of their AI projects in production is set to double within six months (Deloitte, 2026). Yet 74% of organisations report that revenue impact from AI remains aspirational rather than realised, and only 34% are genuinely reimagining core business processes rather than bolting AI onto existing workflows (Deloitte, 2026). The gap between investment and outcome is not a technology gap. It is an architecture, governance, and organisational readiness gap — one that fully on-prem agentic platforms, built on open reasoning models, are uniquely positioned to close.
The release of NVIDIA Nemotron 3 Super in 2025 is the model-level catalyst for this shift. But the more consequential story is what its arrival represents for the enterprise: the moment at which running a sovereign, self-hosted, multi-agent AI automation platform stopped being a calculated trade-off and became a strategic advantage.
What the Data Shows
The Model: Why Nemotron 3 Super Changes the On-Prem Calculus
NVIDIA Nemotron 3 Super is a 120B total-parameter, 12B active-parameter model built on a hybrid Mamba-2 sequence modeling, Transformer attention, and Mixture-of-Experts (MoE) routing architecture (NVIDIA Developer Blog, 2025). This architecture is not a performance curiosity — it is an engineering response to two specific structural barriers that have historically made on-premises enterprise AI deployment economically unworkable at scale.
Barrier 1: Context explosion. Multi-agent systems generate up to 15 times more tokens than standard chat interactions (NVIDIA Developer Blog, 2025). In a typical agentic workflow — where a planning agent, a retrieval agent, an execution agent, and a validation agent must share context across a long task chain — token volume compounds rapidly. Without a model purpose-built for long-context reasoning, agents lose alignment with original objectives. Nemotron 3 Super addresses this with a native 1-million-token context window, the largest available in an open-weight enterprise model.
Barrier 2: The "thinking tax." Extended reasoning models that use chain-of-thought processing incur significant compute overhead — a "thinking tax" — that compounds across every agent in a multi-agent system. For organisations running on-prem GPU clusters, this directly translates to hardware utilisation ceilings and throughput bottlenecks. Nemotron 3 Super's sparse MoE activation means only 12B of its 120B parameters are active during any given inference pass, delivering 4x improved memory and compute efficiency compared to pure transformer architectures, and more than 5x throughput improvement over its predecessor (NVIDIA Developer Blog, 2025).
🔴 Important
The 5x throughput improvement is not relative to cloud models — it is relative to the previous Nemotron Super generation. For on-prem GPU cluster operators, this means existing hardware investments yield dramatically higher agent-per-GPU ratios without additional capital expenditure.
The benchmark evidence is equally striking. Automation Anywhere evaluated Nemotron 3 Super across more than 150 enterprise agent scenarios spanning insurance claims processing, financial operations, healthcare compliance, supply chain management, procurement, and customer retention. The result: Nemotron 3 Super outperformed "almost all large closed-source models" evaluated and was identified as "the best-performing open-source model we have evaluated at its size" (Automation Anywhere, 2025). On PinchBench — a benchmark measuring LLM performance as the reasoning core of an agentic system — the model scores 85.6%, described as "the best open model in its class" (NVIDIA Developer Blog, 2025).
For hardware planning purposes, infrastructure teams should note that running Nemotron 3 Super at 4-bit quantisation requires approximately 64GB–72GB of RAM or VRAM, while 8-bit deployment requires 128GB (Unsloth, 2025). Native NVFP4 pretraining enables 4x faster inference on NVIDIA B200 hardware compared to FP8 on H100 — a meaningful consideration for organisations planning new data centre investments (NVIDIA Developer Blog, 2025).
The Infrastructure: A Converging Reference Architecture
| Component | Specification |
|---|---|
| Primary reasoning model | NVIDIA Nemotron 3 Super (120B/12B active, open weights) |
| Context window | 1M tokens (native) |
| Hardware tier | NVIDIA H100/B200 GPU cluster; Dell AI Factory or equivalent |
| Operating system | Red Hat Enterprise Linux with KVM; VMware vSphere; Bare Metal |
| Orchestration layer | NVIDIA AI Enterprise 5.x (supports OpenShift, vSphere, Bare Metal) |
| Agent framework | Automation Anywhere, EY.ai Agentic Framework, or custom NeMo-based |
| RAG/retrieval | Vector database (pgvector, Milvus, Weaviate, or FAISS on-prem) |
| Vision-language extension | Nemotron Nano VL (for legacy UI and thick-client automation) |
| Governance layer | Agent oversight framework (Deloitte/EY reference models) |
NVIDIA AI Enterprise supports deployment across VMware vSphere, Red Hat Enterprise Linux with KVM, OpenShift on VMware vSphere, Bare Metal Servers, and OpenShift on Bare Metal — providing broad infrastructure compatibility for organisations with heterogeneous on-prem environments (NVIDIA AI Enterprise Deployment Guide, 2025).
The post-training methodology behind Nemotron 3 Super is also relevant to enterprise customisation planning. The model was post-trained with reinforcement learning across 21 environment configurations using more than 1.2 million environment rollouts (NVIDIA Developer Blog, 2025), and is released with fully open weights, datasets, and training recipes — enabling fine-tuning on proprietary enterprise datasets without cloud API exposure.
The Business Context: Productivity Gains Are Real; Revenue Impact Is Not Yet
| Metric | Value | Source |
|---|---|---|
| Organisations reporting productivity/efficiency gains from AI | 66% | Deloitte, 2026 |
| Organisations reporting revenue growth already achieved | 20% | Deloitte, 2026 |
| Organisations genuinely reimagining core processes with AI | 34% | Deloitte, 2026 |
| Organisations using AI at surface level with little process change | 37% | Deloitte, 2026 |
| Organisations with mature governance for autonomous agents | 20% | Deloitte, 2026 |
| Worker AI access growth in 2025 | +50% | Deloitte, 2026 |
| Companies doubling production AI projects within 6 months | Projected 2x | Deloitte, 2026 |
How Leading Organisations Are Responding
Accenture: Pre-Configured Vertical Agents as the Delivery Model
Rather than building bespoke on-premises enterprise AI automation platforms for each client, Accenture has made a decisive architectural bet: standardise the infrastructure layer and compete on vertical agent depth. In collaboration with Dell Technologies and NVIDIA, Accenture launched AI Refinery for on-premises deployment at Dell Technologies World in May 2025, targeting regulated industries with explicit data sovereignty requirements (Accenture newsroom, 2025).
The strategic implication is significant. Accenture is developing more than 50 industry-specific AI agent solutions leveraging NVIDIA reasoning models — with a stated goal of 100 by year-end 2025 — with initial coverage across telecommunications, financial services, and insurance (Accenture newsroom, 2025). As Lan Guan, Accenture's Chief AI Officer, articulated: "Together, we're empowering organisations to accelerate reinvention and unlock new value from data while future-proofing their investments."
This signals a shift in the enterprise AI delivery model: organisations that attempt fully custom agent builds from scratch will face time-to-value disadvantages versus those adopting pre-configured vertical frameworks on a common on-prem infrastructure stack.
EY: The Largest Disclosed Internal Agentic Deployment
EY's approach provides the most concrete public evidence of what scaled on-premises agentic AI looks like in practice. EY.ai enterprise private — powered by the Dell AI Factory with NVIDIA and built on the EY.ai Agentic Framework — is the firm's fully on-premises deployment model (EY newsroom, May 2025). Its internal deployment targets surpassing 3 million tax compliance outcomes and redefining 30 million tax processes annually through 150 AI agents supporting 80,000 professionals (EY newsroom, March 2025).
This is the largest publicly disclosed internal agentic AI deployment figure among the Big Four consulting firms. Critically, EY's choice to deploy on-premises — rather than leverage cloud AI APIs — is explicitly driven by the data sensitivity of tax, risk, and finance workloads for global enterprise clients. The data sovereignty argument is not theoretical for EY; it is the operational reality of managing confidential client financial data across multiple regulatory jurisdictions simultaneously.
Automation Anywhere: Extending On-Prem Automation into Legacy Environments
Automation Anywhere's evaluation of Nemotron 3 Super goes beyond text-based agent workflows. The firm's assessment includes Nemotron vision-language models (Nemotron Nano VL) for UI automation in thick-client and legacy systems — a capability that has significant implications for organisations where traditional Robotic Process Automation (RPA) has hit its limits (Automation Anywhere, 2025).
Many enterprises operate core systems — mainframe interfaces, legacy ERPs, bespoke line-of-business applications — where API-based integration is unavailable and screen-scraping RPA is fragile. Vision-language models that can interpret and act on legacy UI elements extend the reach of an on-premises enterprise AI automation platform into environments that have resisted automation for years. This "second wave" of on-prem capability is substantially underreported in mainstream agentic AI coverage.
💡 Tip
Organisations planning on-prem AI automation deployments should map their legacy system landscape before selecting an agent framework. If thick-client or mainframe interfaces are in scope, prioritise platforms that have validated vision-language model integration for UI automation, not solely API-based agent orchestration.
The Hidden Risk: What Most Teams Get Wrong
The dominant assumption in enterprise AI planning conversations is that the primary risk in on-premises deployment is technical: insufficient compute, model quality gaps versus cloud APIs, or integration complexity. This assumption is wrong — and it is causing organisations to invest in the wrong places.
Deloitte's 2026 State of AI in the Enterprise survey identifies the actual risk with precision: only 1 in 5 companies has a mature governance model for autonomous agents, despite agentic AI usage being "poised to rise sharply in the next two years" (Deloitte, 2026). Jaishiv Prakash, Director Analyst at Gartner, frames the issue directly: "The success of agentic systems will not just depend on model capability but on the overall system architecture, including orchestration, data integration, context management, and governance" (InfoWorld/Gartner, 2025).
⚠️ Warning
Deploying a production-grade on-premises enterprise AI automation platform without mature agent governance infrastructure is not a moderate risk — it is an operational liability. Autonomous agents that make consequential decisions (claims approvals, financial transactions, compliance determinations) without appropriate oversight frameworks expose organisations to regulatory, reputational, and financial risk that scales with deployment breadth.
The specific governance gaps most commonly observed in enterprise agentic deployments include:
- Absence of agent action logging. Many initial deployments lack comprehensive, immutable audit trails of agent decisions and the reasoning chains that produced them — a non-negotiable requirement in regulated industries.
- No human-in-the-loop escalation protocols. Agents operating beyond defined confidence thresholds or encountering novel scenarios require defined escalation pathways to human reviewers. These pathways are frequently absent in first-generation deployments.
- Model drift monitoring gaps. On-prem models, unlike cloud APIs, require active monitoring for performance degradation over time, particularly after fine-tuning on evolving enterprise datasets.
- Scope creep in agent permissions. Multi-agent systems with broad tool-calling permissions tend to expand their operational footprint in ways that were not anticipated at design time.
The governance deficit is compounded by a misunderstanding of what "on-prem" actually achieves. Organisations frequently conflate data residency (data never leaves your infrastructure) with data sovereignty (you have full legal and operational control over data processing and decisions). Achieving the former through on-prem deployment does not automatically achieve the latter — it requires governance architecture that is explicitly designed and enforced.
📘 Note
Data sovereignty is a demand-side forcing function for on-premises enterprise AI, not merely a compliance checkbox. Financial services, healthcare, and insurance organisations subject to GDPR, HIPAA, or sector-specific data localisation laws are not choosing on-prem for performance reasons — they are choosing it because cloud-based AI processing of regulated data may not be legally permissible in their jurisdictions.
A Framework for Moving Forward: The Five Pillars of On-Prem Agentic AI Readiness
Organisations evaluating or accelerating deployment of a fully on-premises enterprise AI automation platform should assess readiness across five distinct pillars. This is not a sequential maturity model — all five must be addressed in parallel, with governance treated as a first-class design constraint rather than a post-deployment retrofit.
| Pillar | Core Question | Minimum Viable Threshold | Leading Practice |
|---|---|---|---|
| 1. Infrastructure Sovereignty | Can you run production agentic workloads on-prem without cloud API dependencies? | NVIDIA AI Enterprise on H100/B200 hardware; 64GB+ VRAM per inference node | Bare metal deployment with NVFP4-optimised B200 for maximum throughput; Dell AI Factory reference architecture |
| 2. Model Selection and Customisation | Does your primary reasoning model meet the performance bar for your specific agentic use cases? | Nemotron 3 Super or equivalent open-weight model validated on representative enterprise scenarios | Fine-tune on proprietary datasets using open training recipes; evaluate vision-language extension for legacy UI coverage |
| 3. Context and Retrieval Architecture | Can your agents maintain coherent context across long, multi-step workflows and retrieve relevant enterprise knowledge at runtime? | On-prem vector database integrated with RAG pipeline; document chunking and embedding infrastructure | 1M-token context model with hybrid retrieval (dense + sparse); domain-specific embedding models fine-tuned on enterprise knowledge corpus |
| 4. Multi-Agent Orchestration | Do you have a defined architecture for how agents communicate, delegate, escalate, and terminate? | Documented agent topology; defined inter-agent communication protocols; scope boundaries per agent | Industry-specific agent frameworks (Automation Anywhere, EY.ai Agentic Framework); tool-calling permission registry |
| 5. Governance and Oversight | Do you have the audit, escalation, monitoring, and compliance infrastructure to operate autonomous agents safely at scale? | Immutable action logging; human-in-the-loop escalation triggers; model drift monitoring | Mature governance model per Deloitte benchmark; agent decision explainability layer; regulatory audit trail aligned to GDPR/HIPAA requirements |
🔴 Important
Pillar 5 — Governance and Oversight — is the most commonly underdeveloped and the most consequential. Organisations should not accelerate Pillars 1–4 faster than their governance infrastructure can support. The 80% of enterprises currently operating without mature agent governance represent a significant risk concentration that will become visible as agentic deployment scales.
The Three Phases of On-Prem Agentic AI Deployment
Phase 1 — Validate (0–90 days): Deploy Nemotron 3 Super on existing GPU infrastructure in a sandboxed environment. Run a representative sample of your target agent scenarios against the Automation Anywhere 150+ scenario framework as a benchmark proxy. Identify your top three highest-value, lowest-risk agentic use cases. Establish baseline governance architecture before any production workloads go live.
Phase 2 — Scale (90–270 days): Integrate on-prem vector databases and RAG pipelines for your primary knowledge domains. Deploy multi-agent orchestration for validated use cases. Implement agent action logging and human-in-the-loop escalation protocols. Begin fine-tuning on proprietary datasets using open training recipes.
Phase 3 — Extend (270+ days): Expand to vision-language agents for legacy UI environments. Adopt pre-configured vertical agent solutions (Accenture AI Refinery, EY.ai Agentic Framework, or equivalent) to accelerate domain coverage. Establish cross-functional agent governance committees. Measure and report agentic AI outcomes against business KPIs, not just technical benchmarks.
What This Means for Your Organisation
The evidence is sufficiently clear to warrant specific, prioritised action — not further evaluation. Here is where your team should direct attention immediately:
1. Resolve your data sovereignty position before your model selection. The question of which LLM to use is downstream of the question of whether your regulated data can legally be processed by a cloud API at all. If you are in financial services, healthcare, insurance, or a government-adjacent sector, conduct a formal data sovereignty assessment before any AI platform procurement decision. This assessment should be led by your General Counsel and Chief Information Security Officer (CISO) jointly, not delegated to the IT architecture team alone.
2. Benchmark Nemotron 3 Super against your specific use cases — not generic leaderboards. Generic LLM benchmarks (MMLU, HumanEval) do not predict agentic task performance in your domain. The Automation Anywhere 150+ enterprise scenario framework is a better reference. Run your top five target agentic workflows against Nemotron 3 Super on your own infrastructure before committing to any platform. Charlie Dai of Forrester notes that the model's hybrid architecture delivers "lower TCO, better utilisation of on-prem or sovereign GPU clusters, and faster agent execution" compared to pure transformers — but verify this claim against your own workload profile (InfoWorld/Forrester, 2025).
3. Treat governance infrastructure as a launch blocker, not a post-launch project. Your organisation is statistically likely to be among the 80% without mature agent governance (Deloitte, 2026). Close this gap before scaling agentic deployments, not after. Assign explicit ownership of agent governance — ideally a named individual with cross-functional authority — and establish the audit trail, escalation, and model monitoring infrastructure before the first production agent goes live.
4. Evaluate pre-configured vertical agent solutions before committing to custom builds. Accenture's model — 100 industry-specific agent solutions by year-end 2025, built on a common NVIDIA/Dell infrastructure stack — represents the direction of enterprise AI delivery. If your sector (telecommunications, financial services, insurance) is among the initial coverage areas, assess whether a pre-configured vertical agent framework accelerates your time-to-value versus a fully custom build. Custom agent development has a role, but it should be reserved for genuinely differentiated use cases, not commodity workflows.
5. Map your legacy system landscape and assess vision-language agent potential. If your organisation has thick-client applications, mainframe interfaces, or legacy ERPs that have resisted RPA integration, Nemotron Nano VL and similar vision-language models represent a materially new capability. Identify the three to five legacy workflows where API-based automation has failed and evaluate whether vision-language agent automation is now viable. This "second wave" of on-prem automation potential is underexplored and may represent your highest-ROI near-term opportunity.
💡 Tip
Organisations that combine RAG-based enterprise knowledge retrieval with vision-language agent UI automation on a common on-premises infrastructure can achieve automation coverage across both structured and unstructured workflows, and across both modern API-accessible systems and legacy environments — without any data leaving their sovereign infrastructure.
Conclusion: The Path Forward
The on-premises enterprise AI automation platform is no longer a compromise — it is a strategic position. Open-weight models have crossed the performance threshold required for production agentic workloads; a reference infrastructure architecture has been validated by the world's largest professional services firms; and data sovereignty regulation is making fully self-hosted AI a legal necessity for an expanding share of enterprise workloads. The organisations that move with urgency on governance infrastructure, not just compute procurement, will be the ones that convert this technical moment into durable competitive advantage. The model is ready. The infrastructure is standardised. The constraint is organisational will.
Sources
- NVIDIA Developer Blog, 2025 — Introducing Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning. https://developer.nvidia.com/blog/introducing-nemotron-3-super-an-open-hybrid-mamba-transformer-moe-for-agentic-reasoning/
- Automation Anywhere, 2025 — Governed On-Prem AI Automation: Delivering a Fully On-Prem Enterprise AI Automation Platform with NVIDIA Nemotron Super. https://www.automationanywhere.com/company/blog/general-technology/delivering-a-fully-on-prem-enterprise-ai-automation-platform-with-nvidia-nemotron-super
- Accenture Newsroom, 2025 — Accenture Collaborates with Dell Technologies and NVIDIA to Accelerate Enterprise AI Transformation with AI Refinery. https://newsroom.accenture.com/news/2025/accenture-collaborates-with-dell-technologies-and-nvidia-to-accelerate-enterprise-ai-transformation-with-ai-refinery
- Accenture Newsroom, 2025 — Accenture Expands AI Refinery and Launches New Industry Agent Solutions to Accelerate Agentic AI Adoption. https://newsroom.accenture.com/news/2025/accenture-expands-ai-refinery-and-launches-new-industry-agent-solutions-to-accelerate-agentic-ai-adoption
- EY Newsroom, March 2025 — EY launching EY.ai Agentic Platform, created with NVIDIA AI, to drive multi-sector transformation starting with tax, risk and finance domains. https://www.ey.com/en_gl/newsroom/2025/03/ey-launching-ey-ai-agentic-platform-created-with-nvidia-ai-to-drive-multi-sector-transformation-starting-with-tax-risk-and-finance-domains
- EY Newsroom, May 2025 — EY announces EY.ai enterprise private, powered by Dell Technologies and NVIDIA accelerated computing to deliver enterprise agentic and physical AI at scale. https://www.ey.com/en_gl/newsroom/2025/05/ey-announces-ey-dot-ai-enterprise-private-powered-by-dell-technologies-and-nvidia-accelerated-computing-to-deliver-enterprise-agentic-and-physical-ai-at-scale
- Deloitte, 2026 — The State of AI in the Enterprise — 2026 AI Report. https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html
- Deloitte, 2025 — Introducing Zora AI™. https://www.deloitte.com/us/en/services/consulting/services/zora-generative-ai-agent.html
- NVIDIA AI Enterprise, 2025 — Deployment Guide — NVIDIA AI Enterprise, Release 5. https://docs.nvidia.com/ai-enterprise/release-5/latest/getting-started/deployment-guide.html
- NVIDIA, 2025 — Foundation Models for Agentic AI — NVIDIA Nemotron. https://www.nvidia.com/en-in/ai-data-science/foundation-models/nemotron/
- Unsloth, 2025 — NVIDIA Nemotron-3-Super: How To Run Guide. https://unsloth.ai/docs/models/nemotron-3-super
- InfoWorld, 2025 — Nvidia launches Nemotron 3 Super to power enterprise AI agents (featuring analysis from Gartner's Jaishiv Prakash and Forrester's Charlie Dai). https://www.infoworld.com/article/4144135/nvidia-launches-nemotron-3-super-to-power-enterprise-ai-agents.html
- FriendliAI, 2025 — Nemotron 3 Super is Live on FriendliAI: multi-agent applications and for specialized agentic AI systems. https://friendli.ai/blog/nvidia-nemotron-3-super
- Google Cloud Blog, 2025 — Now Shipping A4X Max, Vertex AI Training and more. https://cloud.google.com/blog/products/compute/now-shipping-a4x-max-vertex-ai-training-and-more
- Google Cloud Blog, 2025 — Google Cloud AI infrastructure at NVIDIA GTC 2026. https://cloud.google.com/blog/products/compute/google-cloud-ai-infrastructure-at-nvidia-gtc-2026/