Topic: multimodal
8 stories found
Today
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
A new study by TokenSwap identifies a significant gap in consistency among multimodal large language models when responding to semantically equivalent inputs across different modalities, highlighting the need for benchmarking and improvement in these models' performance. This issue matters because it affects the reliability and usability of MLLMs in practical applications where cross-modal input consistency is crucial.
Friday, July 31, 2026
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
A new benchmark called AHA-Memes has been developed to better understand hate in Arabic memes, addressing the growing issue of multimodal online harm where hostile intent is conveyed through a combination of images, text, and cultural references. This matters because it helps advance the detection and analysis of hateful content in underrepresented languages like Arabic.
Thursday, July 30, 2026
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
The study explores how large language models associate gender with musical instruments, revealing potential biases that could reinforce societal stereotypes. This matters because such biases in widely-used AI systems can inadvertently promote gender stereotypes.
Wednesday, July 29, 2026
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
MyoCardBench is a new benchmark for evaluating large language models in realistic cardiovascular care scenarios, addressing limitations of existing benchmarks which often focus on isolated tasks or knowledge rather than longitudinal, multimodal, and safety-critical clinical workflows.
Monday, July 27, 2026
Friday, July 24, 2026
šæ That's all for now. Come back tomorrow.
8 of 8 items shown. Sources: 75 days indexed.