A new study by TokenSwap identifies a significant gap in consistency among multimodal large language models when responding to semantically equivalent inputs across different modalities, highlighting the need for benchmarking and improvement in these models' performance. This issue matters because it affects the reliability and usability of MLLMs in practical applications where cross-modal input consistency is crucial.