parshiv@iitr: ~/portfolio

$ whoami

Parshiv Kapoor

$ echo $ROLE

AI / ML Researcher & Builder

type help to explore  ·  contact to reach me  ·  or press ⌘K

visitor@parshiv:~$
Parshiv Kapoor

Pre-final year B.Tech @ IIT Roorkee

I turn research into reliable, practical AI systems, working across vision-language reasoning, medical imaging, and AI agents.

Download CV GitHub LinkedIn Email

About

I'm a pre-final year undergraduate at IIT Roorkee, focused on building AI systems that are robust, interpretable, and useful in the real world. My work spans medical imaging, reasoning in vision-language models through scene-graph generation, and autonomous AI agents.

I've co-authored papers at AAAI (Student Abstract Track) and ICVGIP 2025, and enjoy taking ideas from a paper to a working, dependable pipeline.

focus:      [ vision-language models, AI agents & agent memory, medical imaging ]
publishing: [ AAAI 2026 (x2, Student Abstract), ICVGIP 2025 ]
building:   [ Warden, FocusBoard, RecallMind ]
now:        exploring the open problems in agent memory

Highlights

SeePhys Challenge @ ICML 2025

Ranked 5th globally on a large-scale vision-language physics-reasoning benchmark (2,000+ multimodal problems, 2,200+ diagrams) with a VLM pipeline that fuses schematic interpretation with natural-language reasoning.

Publications

  • ICVGIP 2025: Tuberculosis bacilli detection from cytopathology images.
  • AAAI 2026 (SAPP): Dual-layer Scene Graph Chain-of-Thoughts for VLM reasoning.
  • AAAI 2026 (SAPP): TWIST, weakly-supervised triplet recognition in surgical video.

Amazon ML Challenge 2025

Ranked 180 of ~10,000 participants on a product price-prediction task using multimodal embeddings.

Focus Areas

Vision-Language Models Medical Imaging AI Agents Scene-Graph Reasoning Robustness & Interpretability Multi-Agent Systems