A Robust Two-Stage Retrieval-Augmented Vision-Language Framework for Knowledge-Intensive Multimodal Reasoning and Alignment. Computational Discovery and Intelligent Systems (CDIS), v. 2, n. 2, p. 42–52, 5 Feb.2026.