Anurag Das

PhD Student at Max Planck Institute for Informatics

Email  /  CV  /  Google Scholar  /  Twitter  /  Github

profile photo

Background

Research Interests

I study how multimodal language models reason across images and views, and how to make them more reliable and efficient. My earlier work focused on semantic segmentation with limited supervision.


Publications
OneScene-Bench: Do Vision-Language Models Integrate Evidence Across Views?
Swathi Suhas, Anurag Das, Bernt Schiele, Jonas Fischer
Under review, 2026

OneScene-Bench tests whether vision-language models combine evidence across views of the same scene. Across counting, consistency, and spatial reasoning tasks, twelve models struggle when cross-view correspondence is required.

More Images, More Problems? A Controlled Analysis of VLM Failure Modes
Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez
ACL Findings, 2026

MIMIC exposes multi-image reasoning failures in large vision-language models. Data generation and attention masking improve cross-image reasoning and achieve state-of-the-art results on multi-image benchmarks.

ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
Mriganka Nath*, Anurag Das*, Jiahao Xie, Bernt Schiele
arXiv, 2026

ClipTTT uses CLIP-guided test-time adaptation on a single image to reduce hallucinations across 15 common corruptions, without changing the base model.

Do Instance Priors Help Weakly Supervised Semantic Segmentation?
Anurag Das*, Anna Kukleva, Xinting Hu, Yuki Asano, Bernt Schiele
Transactions of Machine Learning Research, 2026

SeSAM adapts SAM to weakly supervised segmentation with instance decomposition and iterative pseudo-label refinement. Paired with semi-supervised learning, it improves segmentation while reducing the need for dense annotations.

MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
Anurag Das, Xinting Hu, Li Jiang, Bernt Schiele
ECCV, 2024

MTA-CLIP aligns masks with text to improve semantic segmentation, outperforming pixel–text alignment.

Improving 2D Feature Representations by 3D-Aware Fine-Tuning
Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang, Jan Eric Lenssen
ECCV, 2024

3D-aware fine-tuning improves 2D model features for semantic segmentation and depth estimation.

Weakly-Supervised Domain Adaptive Semantic Segmentation with Prototypical Contrastive Learning
Anurag Das, Yongqin Xian, Dengxin Dai, Bernt Schiele
CVPR, 2023

A unified framework uses image-level, point, and coarse target-domain labels to narrow the gap between unsupervised domain adaptation and supervised segmentation.

Urban Scene Semantic Segmentation with Low-Cost Coarse Annotation
Anurag Das, Yongqin Xian, Yang He, Zeynep Akata, Bernt Schiele
WACV, 2023

Coarse annotations and synthetic data deliver competitive urban-scene segmentation at a fraction of the annotation budget required for fine labels.

(SP)2Net for Generalized Zero-Label Semantic Segmentation
Anurag Das, Yongqin Xian, Yang He, Zeynep Akata, Bernt Schiele
GCPR, 2021 [ Best Master Thesis Award in Germany]

Superpixel pooling supplies class-agnostic segmentation cues, improving generalization to unseen classes in zero-label semantic segmentation.


Thesis
Master Thesis Generalised Zero-shot semantic segmentation with superpixels
Saarland University, 2021

Academics

  • Reviewer: CVPR, ICCV, ECCV, WACV, PAMI, IJCV, TMLR, NeurIPS, ICLR, IEEE-TCVST
  • 2021: DAGM Best Master Thesis Award 2021 in Germany
  • 2021: Saarland Stipendium funded by DAAD
  • 2019: SaarbrĂźcken Graduate School of Computer Science Scholarship
  • 2017: Institute Merit Scholarship (IIIT Allahabad)
  • 2011: Qualified for Indian National Mathematical Olympiad (INMO)



Template from Jon Barron. Big thanks!