India may not need to win every frontier-model benchmark to build a powerful AI ecosystem. The larger opportunity could lie in making intelligence affordable, multilingual, reliable and deeply integrated into enterprise workflows. Recent developments around Sarvam AI offer a glimpse of what an Indian deployment advantage could look like.

For much of the past three years, the global artificial intelligence race has been framed around a familiar set of questions. Which company has the most capable model? Which system performs best on reasoning, coding or mathematics benchmarks? Who has access to the largest amount of compute, and how quickly can the frontier move forward?

Those questions matter. General-purpose AI models are becoming substantially more capable, and the latest generations from OpenAI, Google, Anthropic and others show that progress in reasoning, software engineering, multimodal understanding and computer use is continuing at remarkable speed.

But this may not fully describe the AI opportunity emerging in India.

India operates under a different set of conditions: enormous population scale, extraordinary linguistic diversity, large differences in digital maturity and a market where the economics of technology matter almost as much as technical capability. In such an environment, the most useful AI system may not always be the one with the highest benchmark score. It may be the one that can perform a particular job reliably, in the language of the user, at a cost low enough to be repeated millions of times.

Recent developments around Sarvam AI provide an interesting illustration of this alternative path. A voice AI deployment built by Mahindra AI on Sarvam’s platform has now crossed one crore calls across 12 Indian languages, covering sales, collections and employee engagement at Mahindra Finance. The scale is significant, but the more important signal is that AI is moving beyond controlled demonstrations and entering recurring enterprise workflows.

That distinction deserves attention because the next phase of AI adoption in India may depend less on model size and more on how effectively intelligence can be converted into dependable systems at scale.

The Last Mile of AI Is Harder Than It Looks

An AI demonstration can look convincing in a controlled environment. A model answers a question, summarizes a document or conducts a short voice conversation, and the underlying capability appears obvious.

Production environments are considerably less forgiving.

A customer may begin a conversation in Hindi and switch to English financial terminology midway through a sentence. Another may speak Tamil with a regional accent while calling from a noisy roadside. Product names, locations and technical terms may be uncommon enough for a general speech-recognition system to misinterpret them. The conversation may then need to update a customer record, trigger another workflow or move seamlessly to a human employee when the situation exceeds the agent’s authority.

None of these conditions is unusual in India. Together, they explain why multilingual AI cannot simply be treated as translation layered over an English-first model. Systems built for India need to handle code-mixing, accent variation, noisy audio, specialist terminology and language switching as normal operating conditions.

Sarvam’s Saaras V4 illustrates some of the engineering being directed at this layer. The model is designed to work across all 22 Indian languages supported by Sarvam, improve language identification, remain robust in noisy and code-mixed audio and extend speech recognition beyond Indian English to wider global English accents.

A further September update introduced keyterm prompting, allowing developers to provide up to 50 names, places, brands or technical terms that the speech-recognition system should pay particular attention to.

On paper, that may appear to be a relatively modest feature compared with the excitement surrounding giant foundation models. In production, however, consistently recognising the name of a financial product, customer, village or company can matter more than another incremental gain on a general reasoning benchmark.

This is one of the important distinctions in the next phase of AI: frontier intelligence and deployed intelligence are not the same thing.

AI Value Emerges From the System Around the Model

A useful enterprise AI deployment is rarely just a model.

The voice agent may conduct the conversation, but another system must decide whom to contact, retrieve the relevant information and determine what action should follow. Once the interaction is complete, its outcome has to reach the appropriate enterprise application so that the next step can occur.

The architecture therefore begins to look less like a chatbot and more like a connected chain of decisions.

This becomes particularly important as the discussion shifts toward agentic AI. Much of the attention around AI agents naturally focuses on whether a model can reason through multiple steps or autonomously use tools. In an enterprise environment, however, those abilities become useful only when they operate alongside business data, permissions, applications, deterministic rules and people.

The model remains important, but it is one layer of a much larger system. The real test is whether the complete architecture can move reliably from information to decision and from decision to action without losing context or creating unnecessary operational risk.

In financial services, that could mean linking a voice interaction with CRM, loan-servicing and collections systems. In manufacturing, the same architecture may connect production genealogy, equipment specifications, operating parameters, quality records and approval workflows. In healthcare, the chain could run from patient interaction to scheduling, records and clinical escalation.

Different industries will require very different implementations, but the underlying principle remains consistent: AI creates greater value when it becomes part of the process rather than an intelligence layer sitting beside it.

The Mahindra Finance example is useful precisely for this reason. Its significance is not simply that an AI system can converse in several languages; it is that the voice layer is being used across actual sales, collections and employee-engagement processes at substantial volume.

Scale Changes the Economics of Intelligence

Once AI moves into production, cost also begins to look different.

For an individual user making a handful of requests, the price difference between a large frontier model and a smaller specialised system may be almost irrelevant. Multiply the same interaction across millions of conversations, however, and inference economics become a fundamental part of the architecture.

Sarvam says its stack is currently processing around 400 million API calls a day, while its voice systems have handled approximately 325 million minutes of voice AI in India over the past year. At such volumes, small changes in inference cost, latency or failure rates can materially alter the economics of a deployment.

This is where India could develop a distinctive approach to enterprise AI.

The largest available model may be justified for difficult research, ambiguous reasoning or high-value decisions. It may be unnecessary for a predictable task such as verifying information, routing an enquiry, extracting structured data or conducting a relatively constrained customer conversation.

A more efficient architecture can therefore distribute work across different forms of intelligence. A specialised speech model handles recognition, a smaller language model manages routine conversation, deterministic software enforces hard business rules, and a more capable model is invoked only when deeper reasoning becomes necessary. Humans remain responsible where judgment or authority makes automation inappropriate.

Sarvam itself has been emphasizing this specialised-model approach. The company has cited a 105-billion-parameter model that, for a particular voice workload, it says operates at roughly one-eleventh the cost of certain GPT Mini or Gemini Flash configurations while performing better on the voice metrics it considers relevant. That is a company-reported comparison rather than an independent benchmark, so it should be interpreted accordingly. The underlying architectural argument, however, is broader: the most capable model is not automatically the most commercially useful one.

This leads to a better economic measure for agentic AI. Instead of looking only at cost per million tokens, enterprises may increasingly need to understand cost per successfully completed workflow.

A cheap model that repeatedly fails, escalates incorrectly or requires continuous human correction can become expensive in practice. A more capable or more specialised system that completes the process with fewer interventions may deliver better economics despite a higher unit inference price.

The business process, not the token, ultimately determines value.

Production AI Needs to Know Its Boundaries

The ability to handle exceptions may become just as important as the ability to automate the normal path.

Sarvam’s recent Voice Agents updates provide a good example. An active call can now be transferred to a human, another telephone number or an organisation’s SIP trunk when the conversation requires escalation or specialist intervention. Its webhook capabilities can also return information such as call transcripts, duration and agent variables to downstream systems after the interaction.

These are not spectacular AI capabilities in the conventional sense. They are nevertheless crucial components of production architecture.

Real processes contain ambiguity. Customers ask unexpected questions. A collections discussion may become sensitive. A particular decision may require formal approval. An AI system may simply lack enough confidence to continue safely.

In those circumstances, a well-designed agent should not improvise indefinitely. It should understand the boundary of its role and transfer control appropriately.

This is why human-in-the-loop AI should not be understood merely as a temporary limitation that will disappear as models become smarter. In many regulated, financial, industrial or high-consequence environments, human intervention can be a deliberate design choice.

The goal is not necessarily maximum autonomy. The better objective is appropriate autonomy.

A system that resolves a large majority of routine interactions independently and routes unusual cases intelligently may produce far more business value than one pursuing complete automation while occasionally making costly mistakes.

Sovereign AI Is More Than Building an Indian LLM

The same reasoning also broadens the discussion around sovereign AI.

The phrase is often reduced to whether India can build its own large foundation models. Domestic model development certainly matters, especially for strategic capability, language coverage and sensitive workloads. But genuine AI capability extends beyond the foundation model.

It includes compute infrastructure, the economics of serving models, control over data, language and speech systems, developer platforms, agent frameworks, enterprise integration and governance.

Recent discussion around Sarvam has similarly framed the required Indian stack as extending across compute, models, platforms, agents and governance rather than stopping at model development.

Seen this way, sovereign AI does not require technological isolation.

An Indian enterprise could use a domestic speech model for local-language interactions, an open-source model for certain internal tasks and a global frontier model when unusually difficult reasoning is justified, while retaining its own orchestration, data controls and business logic.

If organisations have credible alternatives across the stack, they can decide where cost, performance, privacy, localisation or strategic control matters most. Dependence decreases not because every component is domestically produced, but because no single external technology becomes impossible to replace.

India Could Compete on the Deployment Curve

The global frontier-model race will continue, and India should participate in it. Better reasoning, stronger multimodal systems, computer use and increasingly capable agents will expand what artificial intelligence can do.

But another race is developing beneath the frontier.

It is the race to make AI dependable across imperfect networks, multiple languages, constrained budgets, legacy enterprise systems and hundreds of millions of people whose interactions with technology differ enormously.

That challenge may play directly to India’s strengths.

The milestone of one crore multilingual AI calls matters, therefore, not simply because the number is large. It provides evidence that artificial intelligence is beginning to cross the difficult boundary between technical capability and operational scale.

The next stage of India’s AI story may consequently be less about reproducing every element of the global model race and more about solving a harder local problem: how do we make intelligence inexpensive enough, reliable enough, multilingual enough and integrated enough to become part of everyday economic activity?

If India can solve that problem, the result may be more consequential than placing another model near the top of a benchmark table. It could create an AI ecosystem built around the realities of large-scale deployment — combining capable models with local languages, affordable inference, enterprise integration and human oversight.

India’s defining AI advantage may ultimately be not how large a model it can build, but how widely, reliably and economically it can put intelligence to work.


Discover more from Poniak Times

Subscribe to get the latest posts sent to your email.