MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation Conference

Al Mazid, MA, Deng, L, Rishe, N. (2026). MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation . 1039-1044. 10.1109/CAI68641.2026.11536347

cited authors

  • Al Mazid, MA; Deng, L; Rishe, N

authors

abstract

  • Clouds remain a major obstacle in optical satellite imaging, limiting accurate environmental and climate analysis. To address the strong spectral variability and the large scale differences among cloud types, we propose MSCloudCAM, a novel multi-scale context adapter network with convolution based cross-attention tailored for multispectral and multi-sensor cloud segmentation. A key contribution of MSCloudCAM is the explicit modeling of multiple complementary multi-scale context extractors. Our formulation uses one extractor's fine-resolution features and the other extractor's global contextual representations enabling dynamic, scale-aware feature selection. Building on this idea, we design a new convolution-based cross attention adaptation module (CAM) that effectively fuses localized, detailed information with broader multi-scale context rather than simply stacking or concatenating context extractors' outputs. Integrated with a hierarchical vision backbone and refined through channel and spatial attention mechanisms, MSCloudCAM achieves strong spectral-spatial discrimination. Experiments on multisensor datasets e.g. CloudSEN12 (Sentinel-2) and L8Biome (Landsat-8), demonstrate that MSCloudCAM achieves superior overall segmentation performance and competitive class-wise accuracy compared to recent state-of-the-art models, while maintaining efficient model complexity, highlighting the effectiveness of the proposed design for large-scale Earth observation.

publication date

  • January 1, 2026

Digital Object Identifier (DOI)

start page

  • 1039

end page

  • 1044