|
I study how multimodal language models reason across images and views, and how to make them more reliable and efficient. My earlier work focused on semantic segmentation with limited supervision.
|
|
OneScene-Bench: Do Vision-Language Models Integrate Evidence Across Views?
Swathi Suhas,
Anurag Das,
Bernt Schiele,
Jonas Fischer
Under review, 2026
OneScene-Bench tests whether vision-language models combine evidence across views of the same scene. Across counting, consistency, and spatial reasoning tasks, twelve models struggle when cross-view correspondence is required.
|
|
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
Anurag Das,
Adrian Bulat,
Alberto Baldrati,
Ioannis Metaxas,
Bernt Schiele,
Georgios Tzimiropoulos,
Brais Martinez
ACL Findings, 2026
MIMIC exposes multi-image reasoning failures in large vision-language models. Data generation and attention masking improve cross-image reasoning and achieve state-of-the-art results on multi-image benchmarks.
|
|
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
Mriganka Nath*,
Anurag Das*,
Jiahao Xie,
Bernt Schiele
arXiv, 2026
ClipTTT uses CLIP-guided test-time adaptation on a single image to reduce hallucinations across 15 common corruptions, without changing the base model.
|
|
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
Anurag Das*,
Anna Kukleva,
Xinting Hu,
Yuki Asano,
Bernt Schiele
Transactions of Machine Learning Research, 2026
SeSAM adapts SAM to weakly supervised segmentation with instance decomposition and iterative pseudo-label refinement. Paired with semi-supervised learning, it improves segmentation while reducing the need for dense annotations.
|
|
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
Anurag Das,
Xinting Hu,
Li Jiang,
Bernt Schiele
ECCV, 2024
MTA-CLIP aligns masks with text to improve semantic segmentation, outperforming pixelâtext alignment.
|
|
Improving 2D Feature Representations by 3D-Aware Fine-Tuning
Yuanwen Yue,
Anurag Das,
Francis Engelmann,
Siyu Tang,
Jan Eric Lenssen
ECCV, 2024
3D-aware fine-tuning improves 2D model features for semantic segmentation and depth estimation.
|
|
Weakly-Supervised Domain Adaptive Semantic Segmentation with Prototypical Contrastive Learning
Anurag Das,
Yongqin Xian,
Dengxin Dai,
Bernt Schiele
CVPR, 2023
A unified framework uses image-level, point, and coarse target-domain labels to narrow the gap between unsupervised domain adaptation and supervised segmentation.
|
|
Urban Scene Semantic Segmentation with Low-Cost Coarse Annotation
Anurag Das,
Yongqin Xian,
Yang He,
Zeynep Akata,
Bernt Schiele
WACV, 2023
Coarse annotations and synthetic data deliver competitive urban-scene segmentation at a fraction of the annotation budget required for fine labels.
|
|
(SP)2Net for Generalized Zero-Label Semantic
Segmentation
Anurag Das,
Yongqin Xian,
Yang He,
Zeynep Akata,
Bernt Schiele
GCPR, 2021 [ Best Master Thesis Award in Germany]
Superpixel pooling supplies class-agnostic segmentation cues, improving generalization to unseen classes in zero-label semantic segmentation.
|
|
-
Reviewer: CVPR, ICCV, ECCV, WACV, PAMI, IJCV, TMLR, NeurIPS, ICLR, IEEE-TCVST
-
2021: DAGM Best Master Thesis Award 2021 in Germany
-
2021: Saarland Stipendium funded by DAAD
-
2019: SaarbrĂźcken Graduate School of Computer Science Scholarship
-
2017: Institute Merit Scholarship (IIIT Allahabad)
-
2011: Qualified for Indian National Mathematical Olympiad (INMO)
|
|