CORA Exposes Thinking-Answer Gap in Multimodal RLVR
CORA identifies a previously underestimated semantic inconsistency in multimodal RLVR and introduces a consistency-oriented alignment method. The findings are promising but require broader validation beyond current benchmarks.











