CV
Arka Mukherjee — Undergraduate researcher in multimodal LLMs and VLM evaluation.
Contact Information
| Name | Arka Mukherjee |
| Professional Title | Undergraduate Researcher |
| arka.mukherjee078@gmail.com |
Professional Summary
Undergraduate researcher focused on multimodal LLM evaluation, reasoning benchmarks, and AI agents. Incoming Research Intern on the AMD AGI team working on LLM inference, Research Fellow at IIT Bhubaneswar (with Dr. Shreya Ghosh), and IUSSTF-Viterbi Summer Research Intern at USC. CS junior at KIIT University (GPA 9.74/10). Published at ICCV 2025, IJCNLP-AACL 2025, and EMNLP 2026 Findings.
Experience
-
- Bengaluru, India
Incoming Research Intern
AMD
AMD AGI Team
- Working on LLM inference.
-
2026 - 2026 Los Angeles, CA
IUSSTF-Viterbi Summer Research Intern
University of Southern California
Advisor: Dr. Maja Mataric
- Built LazySloth, a fast video understanding method that improved retrieval efficiency by up to 8.3x over existing pipelines such as VideoLucy and WorldMM, while retaining task accuracy. (Submitted to AAAI 2027)
- Contributed to lab infrastructure on LLM inference.
-
2026 - 2026 Remote
Research Intern
Carnegie Mellon University
Advisor: Dr. Min Xu
- Applied GRPO, DPO, and PPO to study RL generalization to codon translation, a low-sampling decoding task (64-token action space).
- Improved VRAM efficiency of the code to iterate GRPO on 800k+ datapoints.
-
2024 - Bhubaneswar, India
Undergraduate Research Fellow (Funded)
IIT Bhubaneswar
Advisor: Dr. Shreya Ghosh
- Developed mmJEE-Eval, a 1,460-problem multimodal STEM reasoning benchmark evaluating 17 VLMs. Discovered metacognitive barriers: VLMs detect 53% of errors but correct only 3.5%. (IJCNLP-AACL 2025 Findings)
- Created the first evaluation framework for VLM cultural competence via multimodal story generation on 5 VLMs. (Oral @ ICCV 2025 ASI Workshop)
- Proposed ICE, a new evaluation paradigm unlike benchmarks and leaderboards that models real-world tasks.
- Engineered CoTDL-mini-SWE-agent, a novel coding harness to improve frontier agentic terminal coding performance.
-
2025 - 2025 Ropar, India
Summer Research Fellow (IASc-INSA-NASI)
IIT Ropar
Advisor: Dr. Sudarshan Iyengar
- Developed EduVLM-Bench for STEM prerequisite detection and evaluated 5 open-source LLMs. Top model (Gemma3 27B) achieved 38.5% accuracy.
Education
-
2023 - 2027 Bengaluru, India
B.Tech CSE
Kalinga Institute of Industrial Technology (KIIT)
Computer Science and Systems Engineering
- Agentic AI, Probability & Statistics, Machine Learning, Algorithms, Data Mining, Human-Computer Interaction
Publications
-
2026 -
2025 Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation
ICCV 2025 Workshop on Artificial Social Intelligence (ASI) - Oral
-
2025
Awards
-
2026 Singapore AI Safety Hub (SASH) FAST Fellow
Singapore AI Safety Hub
Selected for the fully funded Frontier AI Security Training fellowship with a 2.8% acceptance rate.
-
2025 KIIT Merit Scholar (Dean's List)
KIIT University
Dean’s List recognition for 6 semesters.
-
2026 Amgen Scholars Program
Amgen Foundation
Selected for the Amgen Scholars Summer Research Program (declined).
-
2026 Aalto Science Institute (AScI) Summer Research Fellow
Aalto University, Finland
Selected for a summer research position in Finland (declined).
-
2025 D&I Subsidy Award ($750)
IJCNLP-AACL 2025
Travel grant for presenting at IJCNLP-AACL 2025.
-
2025 NeurIPS 2025 DCVLR Challenge - #6/59
NeurIPS 2025
Team Blackwell ranked 6th out of 59 teams (top 10th percentile).