The advances in modern sensor technology have greatly expanded the use of remote sensing images in scientific research and in many other life activities of humankind [1-3]. However, in practice, there are always some trade-offs in the design of remote sensing instruments due to technical and budget limitations. Satellites which image a wider swath width do have a shorter revisiting period, but this usually decreases the spatial resolution of observed images, and vice versa [4]. Currently, it is not easy to acquire images that have both high spatial and high temporal resolution [5, 6]. For example, the widely used Landsat images have enabled a 30-m spatial resolution in the visible and infrared spectral bands since the Landsat 4 was launched in 1982 [7]. The latest Landsat 8 maintains a 30-m resolution in most spectral bands, with a long revisiting period of 16 days (the same as Landsat 4, for data continuity purposes) [7]. Conversely, the MODerate Resolution Imaging Spectroradiometer (MODIS) instruments acquire data only at spatial resolution of 250 to 1000 m in multiple spectral bands, but MODIS provides daily coverage of most parts of our planet [8]. For the purpose of long-time series analysis of high spatial resolution imagery (e.g., vegetation-index-based monitoring of crop condition and anomalies at field scale [9, 10], as well as water resource assessment [11]), a single high spatial resolution data source usually cannot meet the requirements of frequent temporal coverage. A number of remote sensing data fusion algorithms have been put forward to address this problem, and research has shown that generating high spatiotemporal data by fusing high spatial resolution images and high temporal resolution images from multiple data sources is a practical approach [12, 13].

In the remote sensing domain, spatiotemporal data fusion refers to a class of techniques that merge two or more data sources which share similar spectral ranges to generate high spatiotemporal time-series data and to derive richer information than a single data source can provide. In most cases, one data source has high temporal but low spatial resolution (HTLS), while another has low temporal but high spatial resolution (LTHS). After years of development, the research field of spatiotemporal data fusion has established certain theories and methods, and some of these methods have been applied in practical geoscience analysis with respectable accuracy [13-15]. As far as we have considered, the existing spatiotemporal fusion algorithms can be classified into three categories: (1) transformation-based; (2) reconstruction-based; and (3) learning-based [16].

The transformation-based methods employ specialized mathematical transforms, such as wavelet transform [17], to transform data from spatial domain to another domain—typically to a frequency or frequency-equivalent domain. Clear, high-frequency components are then extracted from the transformed LTHS images and are merged with HTLS images using elaborately designed fusion rules. The theoretical basis of this approach is that images in different spaces reveal different types of features, which allows a well-designed algorithm to catch the desired features from specific spaces via appropriate transformations.

The reconstruction-based methods generate composite images from weighted sums of spectrally similar neighboring pixels in the HTLS and LTHS image pairs [16]. At present, most of the spatiotemporal fusion algorithms fall into this category. The reconstruction-based methods can be further subdivided into two major groups: one is based on the ground coverage changes of different temporal images, and another is based on the components of mixed ground material end member fractions. In the first case, a hypothetical relation is established based on the difference or ratio deviation between the HTLS and LTHS image pair at the prediction time and a second image pair at a given reference time. Then, a moving window is employed to scan similar neighboring pixels locally and determine the weights. The final composite image is generated by a weighted sum of neighboring pixels in the moving window combined with the hypothetical relation. A typical example is the spatial and temporal adaptive reflectance fusion model (STARFM) [12]. It uses the differences between HTLS and LTHS to establish a relation, and searches similar neighboring pixels by spectral difference, temporal difference, and location distance. Inspired by STARFM, some other enhanced or improved fusion models have been proposed, such as the spatial and temporal adaptive algorithm for mapping reflectance change (STAARCH) [18], enhanced STARFM (ESTARFM) [5], and other STARFM-based models [19]. In general, the main differences among algorithms of this type are in the designs of HTLS and LTHS relations and in the rules used to determine weights.

Spatiotemporal Data Fusion in Remote Sensing: A Review of Techniques and Applications

原文地址: https://www.cveoy.top/t/topic/famr 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录