publications

2026

  1. Preprint
    IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
    Varun Gumma, Navonil Majumder, Soumitra Sinhahajari, and Soujanya Poria
    2026
  2. ACL
    UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
    Pranjal A Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali, Manan Uppadhyay, Deepthi Sudharsan, and Sunayana Sitaram
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2026
  3. ICLR
    OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
    Jingdi Lei*, Varun Gumma*, Rishabh Bhardwaj*, Seok Min Lim, Chuan Li, Amir Zadeh, and Soujanya Poria
    In The Fourteenth International Conference on Learning Representations, 2026

2025

  1. Preprint
    HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
    Varun Gumma, Ananditha Raghunath, Mohit Jain, and Sunayana Sitaram
    2025
  2. NAACL
    Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
    Varun Gumma*, Pranjal A Chitale*, and Kalika Bali
    In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Apr 2025
  3. AfricaNLP
    Beyond Metrics: Evaluating LLMs Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
    Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Vishrav Chaudhary, Keshet Ronen, Kalika Bali, and Jacki O’Neill
    In Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025), Jul 2025

2024

  1. EvalEval
    Contamination Report for Multilingual Benchmarks
    Sanchit Ahuja*, Varun Gumma*, and Sunayana Sitaram
    2024
  2. EMNLP
    PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
    Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  3. FAccT
    Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
    Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, Ujwal Gadiraju, Aditya Vashistha, Vivek Seshadri, and Kalika Bali
    In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro, Brazil, 2024
  4. NAACL
    METAL: Towards Multilingual Meta-Evaluation
    Rishav Hada*, Varun Gumma*, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram
    In Findings of the Association for Computational Linguistics: NAACL 2024, Jun 2024
  5. NAACL
    MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
    Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram
    In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Jun 2024
  6. EACL
    Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
    Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram
    In Findings of the Association for Computational Linguistics: EACL 2024, Mar 2024
  7. EACL
    MAFIA: Multi-Adapter Fused Inclusive Language Models
    Prachi Jain*, Ashutosh Sathe*, Varun Gumma, Kabir Ahuja, and Sunayana Sitaram
    In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Mar 2024
  8. ComputEL
    MunTTS: A Text-to-Speech System for Mundari
    Varun Gumma, Rishav Hada, Aditya Yadavalli, Pamir Gogoi, Ishani Mondal, Vivek Seshadri, and Kalika Bali
    In Proceedings of the Seventh Workshop on the Use of Computational Methods in the Study of Endangered Languages, Mar 2024

2023

  1. TMLR
    IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages
    Jay Gala*, Pranjal A Chitale*, A K Raghavan, Varun Gumma, Sumanth Doddapaneni, Aswanth Kumar M, Janki Atul Nawale, Anupama Sujatha, Ratish Puduppully, Vivek Raghavan, Pratyush Kumar, Mitesh M Khapra, Raj Dabre, and Anoop Kunchukuttan
    Transactions on Machine Learning Research, 2023
  2. EAMT
    An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models
    Varun Gumma, Raj Dabre, and Pratyush Kumar
    In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, Jun 2023

2022

  1. SECRYPT
    PAMMELA: Policy Administration Methodology using Machine Learning
    Varun Gumma, Barsha Mitra, Soumyadeep Dey, Pratik Patel*, Sourabh Suman*, Saptarshi Das, and Jaideep Vaidya
    In Proceedings of the 19th International Conference on Security and Cryptography, 2022