2 Sources
[1]
Gnani AI launches Artha sovereign AI stack with 30-billion-parameter Evon 3.3
Gnani AI launched Artha, a sovereign AI stack for Indian companies and public institutions. This stack features the Evon 3.3 language model and the Plexus agentic platform. Evon 3.3 is an open-weights model trained on Indic languages and domain-specific data. The Artha stack allows organisations to run AI models within their own infrastructure. Voice AI company Gnani AI on Friday (August 28) launched Artha, a sovereign artificial intelligence stack built around its 30-billion-parameter Evon 3.3 language model and an agentic AI platform called Plexus. The stack, aimed at Indian companies and public institutions, was unveiled in New Delhi by Vice President CP Radhakrishnan. According to the company, Artha is designed to allow organisations to run AI models and applications within their own infrastructure, particularly relevant for banks, insurers and government departments that handle sensitive information and face regulatory requirements governing where data is stored and processed. Gnani added that Evon 3.3 is being released as an open-weights model, while Plexus will be offered to enterprise customers. Also Read: Business outcomes, not AI models, will decide enterprise deals: Gnani.ai CEO Ganesh Gopalan Built on Nvidia's Nemotron Gnani developed Evon 3.3 by continually pre-training Nvidia's Nemotron model on its Indic-language and domain-specific data. This was followed by post-training and reinforcement learning. "We have pre-trained this with our own data, with our own tokens. It is also post-trained with reinforcement learning," Gnani cofounder and chief executive Ganesh Gopalan told The Economic Times Digital. "If you compare Evon with the standard Nemotron model, you will see the difference in its Indic-language capabilities, performance, token efficiency and cost. It is significantly more efficient for Indian languages than generic global models in terms of tokens consumed and accuracy on benchmarks such as MILU. You do not see those capabilities in the standard Nemotron model, but you see them in Evon," he added. The training corpus contained more than 2 trillion tokens across 11 Indian languages, according to Gnani cofounder and chief product and engineering officer Bharath Shankar. "We have also optimised Evon for tool calling, language understanding, reasoning, speed and cost," Shankar said. Evon 3.3 uses a mixture-of-experts architecture. Although the model has 30 billion parameters, only about 3.5 billion are activated for a given task. Shankar said this allows it to offer reasoning and language-understanding capabilities while consuming less computing power than larger models. Gnani claims that Evon 3.3 outperforms Sarvam's 30-billion-parameter model and 105-billion-parameter model in 10 of 11 languages on MILU, a benchmark used to evaluate Indic-language understanding. "Our model has 30 billion parameters but it outperformed the 105-billion-parameter model in 10 of the 11 languages on which it was trained," Shankar said. (Source: Gnani.ai) The results are based on Gnani's internal testing across about 40-45 benchmarks. The company also tested Evon 3.3 on MMLU and MMLU-Pro, widely used benchmarks for evaluating reasoning and language understanding. The development process was divided into data cleaning, continual pre-training and post-training, with Gnani using about 1,500 Nvidia GPUs across these stages, Shankar said. Cutting Indian-language token costs Gnani has also rebuilt the model's tokenizer, the component that breaks text into units a model can process, to handle Indian scripts more efficiently. The company claimed Evon 3.3 requires about 20% fewer tokens per Indian-language word than the tokenizer used by the GPT-5 family and less than half the number used by byte-level tokenizers in models such as DeepSeek, Llama and Qwen. Also Read: From 100,000 calls a month to 5 million a day: Inside Vobiz's AI telephony bet "We evaluate the tokenizer using a measure called token fertility, which looks at how many tokens a model needs to represent text in a particular language. In practice, this means the model consumes fewer tokens to process and generate the same amount of content. It can understand the input language more efficiently and does not need as many tokens to represent it," Shankar said. Reducing the number of tokens required to process text can lower inference costs and latency, particularly when organisations are handling large volumes of Indian-language documents or conversations. Gnani claims that running Evon 3.3 could cost between one-third and one-fifth as much as comparable OpenAI models. "Our claim of roughly 39% to 40% greater efficiency is based on the tokenizer's performance. Fewer tokens directly translate into lower costs and reduced latency, making responses faster by a similar margin," Shankar said. According to him, Evon 3.3 can run on hardware such as Nvidia's RTX 6000 Pro or L40S and does not require high-end accelerators such as the B200 or H200, although performance will vary depending on the workload and deployment. The model weights are available by request on Hugging Face under an Apache 2.0 licence. This will allow enterprises to develop derivative models and deploy them in their own data centres or virtual private clouds. "We are building purpose-built models in India, initially with a focus on Indian languages, but we will extend them to other languages as well. For us, the sovereign story is about building in India for the world, rather than building only for consumption in India," Shankar said. From models to enterprise agents Plexus forms the second part of the Artha stack. The platform allows enterprises to combine multiple AI agents into workflows, specify when humans must intervene and monitor the agents' actions. Also Read: OpenAI-Hugging Face incident exposes cybersecurity's 'human-speed' problem It can work with different underlying models, including Evon 3.3, and connect AI agents with documents, enterprise software and conversations, the company said. Potential applications include processing multilingual loan documents, reconciling transactions across bank statements and core banking systems, resolving citizen grievances, underwriting insurance policies and servicing loans, Gopalan said. Enterprise fine-tuning could become one revenue stream for the open-weights model, alongside Gnani's agentic platform and applications. "There are multiple ways in which the commercial angle of launching these open source, open weight models will evolve," Gopalan said.
[2]
Hon'ble Vice-President of India unveils Gnani Artha - the next frontier of Sovereign AI
Built on Evon v3.3, Gnani AI's open-weights model for 11 Indian languages, and Plexus, its agentic AI platform, Gnani Artha brings sovereign frontier-grade intelligence, AI economics that work in India, and real-world impact at scale into a single stack. Shri C. P. Radhakrishnan, the Honourable Vice President of India, today unveiled Gnani Artha, an end-to-end sovereign AI stack for Indian enterprises and public institutions. Gnani Artha is built on two key components: Gnani Evon v3.3 - a 30-billion-parameter open-weights model trained natively across 11 Indian languages - and Gnani Plexus, the company's agentic AI platform. Speaking on the occasion, the Hon'ble Vice-President of India said, "Bharat is making significant strides in AI. India's approach focuses on making AI open, affordable, and accessible, ensuring that innovation uplifts society as a whole. I congratulate Gnani AI on this significant milestone and I am pleased to launch Gnani Artha, which brings together the power of Evon, a large language model, and Plexus, a platform that connects this intelligence with real-world work and institutions. Both these platforms reflect the growing strength of India's technology ecosystem. This initiative shows that our engineers have the capability not only to use frontier technologies, but also to build them. May this initiative continue to building a stronger, more self-reliant, and technologically empowered India." Consider a loan file that arrives at an Indian lender's office: an application form filled in Marathi, six months of bank statements, GST filings, and photographs of multiple identity documents. Reading it takes judgement, not transcription: cross-checking declared income against actual turnover and finding the inconsistencies that matter. Every lender in India handles thousands of these a day. Almost none can do it with AI that reasons well enough in the applicant's own language, at a cost that is sustainable at that scale, and without the file ever leaving the lender's own systems. Handling that file takes three things at the same time: intelligence that reasons in the applicant's own language, economics that survive thousands of files a day, and a way to put both into production inside the lender's own systems. Indian institutions have had these in fragments. Gnani Artha brings them together as one stack. GNANI EVON V3.3 - SOVEREIGN LLM THAT CUTS THE LANGUAGE TAX ON INDIAN AI Evon v3.3 is a 30-billion-parameter model with roughly 3.5 billion parameters active on any given token. On MILU, a widely used Indian-language benchmark spanning eleven languages and dozens of academic and professional subjects, Evon v3.3 outperforms a 105-billion-parameter Indic model on ten of eleven languages and a similarly sized 30B model on all eleven - and achieves parity with a similar-size, hosted global frontier model. Because Evon v3.3 ships as open weights that run on a single node, it can be deployed entirely inside an institution's own data centre or virtual private cloud. This lets banks, insurers and government bodies meet DPDP, RBI and IRDAI residency requirements without customer data ever leaving their own infrastructure. To make the economics work, Gnani AI rebuilt the model's tokenizer - the layer that breaks language into the units a model processes - for Indian scripts. Evon v3.3 needs roughly 20% fewer tokens per Indian-language word than the tokenizer used by the GPT-5 family, and less than half of what byte-level tokenizers such as DeepSeek, Llama and Qwen require. Fewer tokens mean lower compute cost, faster responses, and more usable context per query. GNANI PLEXUS - TURNING SOVEREIGN INTELLIGENCE INTO REAL-WORLD IMPACT Plexus is Gnani AI's agentic AI platform. Each agent is a discrete, identity-bearing unit - closer to an employee than a script - combined with other agents into workflows built around a defined outcome. For example, in grievance resolution by government agencies, a multilingual AI agent captures a citizen's complaint. A reasoning agent then detects patterns across recent complaints - a spike in one district, one recurring failure - files a single consolidated ticket with the department that owns it, and closes the loop with the citizen once it is resolved. In bank reconciliation, an AI agent matches millions of line items across bank statements, core-banking ledgers and payment-switch logs. It clears the clean matches, drafts a root-cause narrative for each genuine exception, and routes it to the team that owns it - automating one of the highest-cost manual functions in Indian banking. These workflows run under an orchestration layer that can be human-in-the-loop or AI-led, with guardrails, observability and audit logging built in rather than added afterwards. Plexus integrates with enterprise and public digital infrastructure, supports tool calling and a choice of underlying models including Gnani Evon v3.3, and works across documents, systems and conversations alike. "Sovereign AI is not about keeping the world out. It is about India having the capability to build for itself - and then for every country that shares its problems, "concluded Gopalan. Evon v3.3. weights are available by request on Hugging Face under an Apache 2.0 licence. Plexus becomes available to enterprise customers.
Share
Copy Link
Gnani AI launched Artha, an end-to-end sovereign AI stack featuring the 30-billion-parameter Evon 3.3 model and Plexus agentic AI platform. Unveiled by Vice-President of India CP Radhakrishnan, the stack enables Indian companies and public institutions to run AI models within their own infrastructure while supporting 11 Indian languages at significantly lower costs than global alternatives.
Gnani AI has launched Artha, an end-to-end sovereign AI stack designed specifically for Indian companies and public institutions.
1
The stack was unveiled in New Delhi by Vice-President of India CP Radhakrishnan on August 28, marking a significant milestone in India's AI capabilities. Built around two core components—the 30-billion-parameter open-weights model Evon 3.3 and the Plexus agentic AI platform—Artha addresses three critical needs: intelligence that reasons in Indian languages, economics that work at scale, and deployment within organizations' own infrastructure.2
Speaking at the launch, the Vice-President of India emphasized that this initiative demonstrates Indian engineers' capability to build frontier technologies, not just use them. The Artha stack is particularly relevant for banks, insurers, and government departments handling sensitive information while facing regulatory requirements around data sovereignty.
1

Source: CXOToday
Evon 3.3 was developed by continually pre-training Nvidia's Nemotron model on Gnani's Indic-language and domain-specific data, followed by post-training and reinforcement learning.
1
The training corpus contained more than 2 trillion tokens across 11 Indian languages, according to Gnani cofounder and chief product and engineering officer Bharath Shankar.Gnani AI claims Evon 3.3 outperforms Sarvam's 30-billion-parameter model and even its 105-billion-parameter model in 10 of 11 languages on the MILU benchmark, which evaluates Indic-language understanding.
1
The results stem from internal testing across approximately 40-45 benchmarks, including widely used standards like MMLU and MMLU-Pro for evaluating reasoning and language understanding.The model employs a mixture-of-experts architecture, with only about 3.5 billion of its 30 billion parameters activated for any given task. This design allows Evon 3.3 to deliver reasoning and language-understanding capabilities while consuming less computing power than larger models.
1
The development process utilized approximately 1,500 Nvidia GPUs across data cleaning, continual pre-training, and post-training stages.Gnani AI rebuilt the model's tokenizer—the component that breaks text into processable units—to handle Indian scripts more efficiently. The company claims Evon 3.3 requires about 20% fewer tokens per Indian-language word than the tokenizer used by the GPT-5 family and less than half the number used by byte-level tokenizers in models such as DeepSeek, Llama, and Qwen.
1
This efficiency translates directly into cost savings. Gnani AI asserts that running Evon 3.3 could cost between one-third and one-fifth as much as comparable OpenAI models, with roughly 39% to 40% greater efficiency based on the custom tokenizer's performance.
1
Fewer tokens directly reduce inference costs and latency, making responses faster by a similar margin—critical when organizations handle large volumes of Indian-language documents or conversations.Evon 3.3 can run on accessible hardware such as Nvidia RTX 6000 Pro or L40S, eliminating the need for high-end accelerators like the B200.
1
Because the model ships as open weights that run on a single node, it can be deployed entirely inside an institution's own data centre or virtual private cloud, enabling banks, insurers, and government bodies to meet DPDP, RBI, and IRDAI residency requirements without customer data ever leaving their infrastructure.2
Related Stories
Plexus serves as Gnani AI's agentic AI platform, where each agent functions as a discrete, identity-bearing unit that can be combined into workflows built around defined outcomes.
2
The platform addresses real-world use cases that have strained Indian institutions for years.In loan processing, consider a file arriving at an Indian lender's office: an application form filled in Marathi, six months of bank statements, GST filings, and photographs of multiple identity documents. Reading this requires judgement—cross-checking declared income against actual turnover and identifying inconsistencies that matter. Every lender in India handles thousands of these daily, but almost none can process them with AI that reasons well enough in the applicant's own language, at sustainable costs, without files leaving their systems.
2
For grievance resolution by government agencies, a multilingual AI agent captures a citizen's complaint. A reasoning agent then detects patterns across recent complaints—a spike in one district, one recurring failure—files a single consolidated ticket with the department that owns it, and closes the loop with the citizen once resolved.
2
In bank reconciliation, an AI agent matches millions of line items across bank statements, core-banking ledgers, and payment-switch logs. It clears clean matches, drafts a root-cause narrative for each genuine exception, and routes it to the team that owns it—automating one of the highest-cost manual functions in Indian banking.
2
These workflows run under an orchestration layer that can be human-in-the-loop or AI-led, with guardrails, observability, and audit logging built in. Plexus integrates with enterprise and public digital infrastructure, supports tool calling and a choice of underlying models including Gnani Evon v3.3, and works across documents, systems, and conversations.
2
While Evon 3.3 is being released as an open-weights model, Plexus will be offered to enterprise customers.1
Summarized by
Navi
13 Aug 2024

24 Apr 2025•Technology

30 Oct 2025•Business and Economy

1
Technology

2
Policy and Regulation

3
Technology
