Nasscom released Hindi–English large-language-model benchmarks on Monday, timed to MeitY’s GenAI sandbox opening, scoring retrieval, summarisation, and code-switch handling so enterprise buyers can compare vendors without relying on English-only demo scripts.
What the benchmark measures
The suite, dubbed IndicShift-HE for internal tracking, spans 1,200 prompt tasks drawn from customer-service logs, RBI circular summaries, and state education portal FAQs—domains where ministries already deploy chatbots. Tasks are split evenly between monolingual Hindi, monolingual English, and mixed sentences with transliterated product names.
Scoring blends automatic metrics—BLEU and BERTScore variants tuned for Devanagari—with human rubrics from Nasscom’s reviewer pool on factual consistency and refusal behaviour when prompts ask for illegal financial advice.
Why Hindi–English first
Central government schemes and large banks still operate bilingually even as states push single-language interfaces. MeitY asked industry bodies for benchmarks that mirror procurement questionnaires due this quarter. Hindi–English code-switching is the highest-volume pair in BHASHINI call-centre audio, according to mission statistics cited in the report.
Nasscom officials said Tamil–English and Telugu–English extensions are underway but need more licensed text; the HE release is the minimum bar for sandbox graduation criteria MeitY circulated to members last week.
Methodology transparency
Dataset IDs map to BHASHINI snapshot versions so teams can reproduce runs inside the sandbox. Nasscom withheld raw logs with customer phone numbers, publishing only paraphrased prompts. Model testers must disclose parameter counts, quantisation, and whether RAG pipelines pulled from live web search—practices that can inflate scores artificially.
Baseline scores for open models hosted on IndiaAI compute are included; several Indian startups volunteered checkpoints under NDA so buyers see relative lifts without exposing weights.
How vendors should use it
Enterprise CIOs get a one-page scorecard template for RFP annexes. Startups are encouraged to publish delta improvements when fine-tuning on sandbox data, not absolute numbers that compare unfairly across hardware. Nasscom warned against “benchmark hacking” via prompt leakage, threatening delisting from its AI directory if gaming is proven.
Global hyperscalers with India regions participated as observers; their India-facing teams can opt in to official runs next quarter when legal clears cross-border logging.
Link to MeitY sandbox
Sandbox graduation now references IndicShift-HE thresholds on harm refusals and factual QA bands. Teams that beat baselines by published margins receive marketplace badges MeitY plans to show on GEM listings. Failure does not block research access but slows procurement endorsements—a nudge aligned with IndiaAI’s responsible deployment memo.
MeitY and Nasscom will co-host reviewer training so human scores stay consistent across cities, reducing the “Bangalore bonus” startups complained about in earlier TTS benchmarks.
Criticism and gaps
Language activists noted dialect coverage is thin for Bhojpuri and Hinglish street slang common in UPI dispute calls. Legal scholars asked whether benchmark text licences allow commercial redistribution; Nasscom said RFP users may cite scores with attribution but not republish prompts.
Some open-source advocates want weights for the automatic metric calculators; Nasscom promised GitHub release after sanitising proprietary cleaner code.
Market impact
IT services firms pitching GenAI modernization can attach Nasscom-backed scorecards instead of bespoke demos that clients cannot audit. Banks may still run their own hold-out tests, but baseline parity reduces weeks of vendor shootouts.
Investors treating Indic LLMs as a category get a crude league table—imperfect, but better than valuing teams on English MMLU scores alone.
What comes next
Quarterly refresh cycles will add election-season misinformation tasks after state polls, plus accessibility prompts for visually impaired users relying on Hindi screen readers. Nasscom is soliciting donated tasks from insurers and telcos under NDAs to broaden domain coverage.
If benchmarks diverge too far from real call-centre outcomes, MeitY can adjust sandbox graduation weights—a feedback loop both sides said they want before ministries sign billion-rupee chatbot contracts.
Documentation for buyers
Nasscom published a buyer’s guide appendix that maps IndicShift-HE metrics to typical RFP clauses, including how to score tie bids when two vendors land within the published margin of error on factual QA tasks. State IT departments in Uttar Pradesh and Maharashtra have already referenced the appendix in draft chatbot tenders circulating this month. Reviewers said they will publish errata if sandbox dataset versions change mid-quarter so scores remain comparable.








