Skip to main navigation Skip to search Skip to main content

VisAffect at MWE-2026 AdMIRe 2: IMMCAN Idiom Multimodal Cross-Attention Network

  • Istanbul Technical University
  • New York University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Citation (Scopus)

Abstract

We address AdMIRe 2.0, a static image ranking task where a sentence containing a potentially idiomatic expression is paired with five image-caption candidates, and the goal is to rank the candidates by semantic compatibility with the intended idiomatic or literal meaning. We propose IMMCAN, which keeps XLM-R and Jina-CLIP-v2 frozen and learns a lightweight two-stage cross-attention fusion, caption-image grounding followed by idiom-to-multimodal conditioning, to predict a compatibility score per candidate. We also evaluate caption-only augmentation via back-translation and synonym substitution, and compare regression and rank-class formulations. On AdMIRe 1.0, text-only achieves higher test top-image accuracy than VLM-grounded modeling. In contrast, on AdMIRe 2.0 zero-shot, adding visual patch grounding improves both accuracy and NDCG indicating better cross-lingual ranking transfer.

Original languageEnglish
Title of host publicationMWE 2026 - 22nd Workshop on Multiword Expressions, Proceedings of the Workshop
EditorsAtul Kr. Ojha, Verginica Barbu Mititelu, Mathieu Constant, Ivelina Stoyanova, A. Seza Dogruoz, Alexandre Rademaker
PublisherAssociation for Computational Linguistics (ACL)
Pages149-153
Number of pages5
ISBN (Electronic)9798891763630
DOIs
Publication statusPublished - 2026
Event22nd Workshop on Multiword Expressions, MWE 2026 - Rabat, Morocco
Duration: 28 Mar 2026 → …

Publication series

NameMWE 2026 - 22nd Workshop on Multiword Expressions, Proceedings of the Workshop

Conference

Conference22nd Workshop on Multiword Expressions, MWE 2026
Country/TerritoryMorocco
CityRabat
Period28/03/26 → …

Bibliographical note

Publisher Copyright:
© 2026 Association for Computational Linguistics.

Fingerprint

Dive into the research topics of 'VisAffect at MWE-2026 AdMIRe 2: IMMCAN Idiom Multimodal Cross-Attention Network'. Together they form a unique fingerprint.

Cite this