Non-local attention improves single-image super-resolution (SISR) by letting every query position aggregate information from all other positions, but the softmax normalisation at its core assigns strictly positive weight to every candidate regardless of relevance, diluting genuinely informative long-range correlations once the candidate set is large. This paper proposes Adaptive Sparse Attention (ASA), a non-local attention mechanism that replaces the fixed softmax support with a per-query threshold ? = ?(s) + ?•?(s) derived from the mean and standard deviation of each query’s own similarity-score distribution, so that only statistically above-average candidates are retained. This further proposes two extensions: SCASA-TT, which compresses the query/key/value projections with a Tensor-Train (TT) factorisation and adds a complementary spatial-coupling sparsity mask, and Quantum ASA, in which a 4-qubit variational quantum circuit (VQC) predicts the threshold coefficient ? per query rather than using a fixed constant. All three variants, together with a non-attentive convolutional baseline, share a common backbone and are evaluated using Y-channel PSNR/SSIM on Set5, Set14, BSD100 and Urban100. Quantum ASA attains the best reconstruction quality of every method on every dataset and scale factor tested, with the largest gains observed on Urban100 and Manga109, the two benchmarks richest in repeated structure, consistent with the premise that non-local self-similarity is a more exploitable prior precisely where such redundancy is abundant. An analytical hardware-cost model projects a 73.7% reduction in multiply-accumulate operations and a 74.8% reduction in memory bandwidth for ASA-sparsified attention relative to dense attention at matched sequence length.
Introduction
1. Introduction
The research proposes an efficient deep-learning framework for Single-Image Super-Resolution (SISR), which improves the resolution and visual quality of low-resolution images. Traditional non-local attention captures long-range similarities between image regions but suffers from two major problems: softmax dilution, where irrelevant image regions receive attention weights, and high computational complexity of O(N2)O(N^2)O(N2).
To address these issues, the study introduces three sparse attention mechanisms: Adaptive Sparse Attention (ASA), SCASA-TT, and Quantum ASA. These methods aim to improve reconstruction quality while reducing unnecessary computations and model complexity.
2. Proposed Methodology
The proposed AdaptiveSRNet framework uses a convolutional backbone with residual blocks, an attention module, and PixelShuffle-based upsampling. Four model variants are evaluated:
Baseline: A convolutional network without attention.
ASA: Uses a query-specific statistical threshold based on the mean and standard deviation of similarity scores to discard irrelevant positions.
SCASA-TT: Combines statistical and spatial-coupling masks with Tensor-Train compression to reduce parameter overhead.
Quantum ASA: Uses a four-qubit variational quantum circuit to predict an adaptive sparsification threshold from similarity statistics.
The models are trained using DIV2K and evaluated on Set5, Set14, BSD100, Urban100, and Manga109 at ×2, ×3, and ×4 scaling factors, using PSNR and SSIM metrics.
3. Key Results
At ×2 scaling, Quantum ASA achieves the highest PSNR and SSIM among the four proposed variants across all evaluated datasets.
Quantum ASA improves PSNR over the baseline by 0.24 dB on BSD100 and 0.56 dB on Urban100.
At ×3 scaling, Quantum ASA achieves a 1.24 dB PSNR improvement on Manga109 compared with the baseline.
At ×4 scaling, it improves PSNR by 0.23–0.62 dB across the four evaluated datasets.
ASA removes approximately 79% of candidate attention positions on average, while SCASA-TT achieves parameter compression with only a 0.36% parameter increase over the baseline.
The study also reports that Quantum ASA learns a threshold coefficient of approximately 0.4356, compared with the fixed ASA threshold coefficient of 0.5.
4. Computational Efficiency
All proposed variants contain fewer than 500,000 trainable parameters and have model sizes below 2.1 MB at FP32 precision. However, measured inference time increases substantially because the current implementation computes dense attention before applying sparsification.
An analytical hardware-cost model projects a 73.7% reduction in multiply-accumulate operations and energy, along with a 74.8% reduction in memory bandwidth. These savings are theoretical projections that require sparse-aware kernels or specialised hardware for practical realisation.
Conclusion
This paper proposed Adaptive Sparse Attention, a statistically-thresholded alternative to dense softmax non-local attention for SISR, together with a Tensor-Train-compressed and spatially-coupled extension (SCASA-TT) and a quantum-circuit-guided variant (Quantum ASA) in which the sparsification threshold is predicted rather than fixed. Across five standard SISR benchmarks and three scale factors, Quantum ASA consistently outperforms a matched-capacity non-attentive baseline and a fixed-threshold statistical variant, with the largest gains on datasets richest in repeated structure. An analytical hardware-cost model indicates substantial multiply-accumulate, bandwidth and energy reduction is achievable from the sparsification pattern learned by these methods; realising this reduction in wall-clock terms through sparse-aware computation is identified as the primary direction for future work.
References
[1] X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 7794–7803.
[2] Y. Mei, Y. Fan, and Y. Zhou, “Image super-resolution with non-local sparse attention,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 3516–3525.
[3] B. Xia, Y. Hang, Y. Tian, W. Yang, Q. Liao, and J. Zhou, “Efficient non-local contrastive attention for image super-resolution,” in Proc. AAAI Conf. Artif. Intell., vol. 36, no. 3, 2022, pp. 2759–2767.
[4] J.-N. Su, M. Gan, G.-Y. Chen, W. Guo, and C. L. P. Chen, “High-similarity-pass attention for single image super-resolution,” IEEE Trans. Image Process., vol. 33, pp. 610–624, 2024.
[5] A. Martins and R. Astudillo, “From softmax to sparsemax: A sparse model of attention and multi-label classification,” in Proc. Int. Conf. Mach. Learn. (ICML), 2016, pp. 1614–1623.
[6] I. V. Oseledets, “Tensor-train decomposition,” SIAM J. Sci. Comput., vol. 33, no. 5, pp. 2295–2317, 2011.
[7] A. Novikov, D. Podoprikhin, A. Osokin, and D. Vetrov, “Tensorizing neural networks,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2015, pp. 442–450.
[8] V. Bergholm et al., “PennyLane: Automatic differentiation of hybrid quantum-classical computations,” arXiv:1811.04968, 2018.
[9] Z. Hui, X. Gao, Y. Yang, and X. Wang, “Lightweight image super-resolution with information multi-distillation network,” in Proc. 27th ACM Int. Conf. Multimedia, 2019, pp. 2024–2032.
[10] J. Liu, J. Tang, and G. Wu, “Residual feature distillation network for lightweight image super-resolution,” in Proc. Eur. Conf. Comput. Vis. Workshops (ECCVW), 2020, pp. 41–55.
[11] S. Anwar and N. Barnes, “Densely residual Laplacian super-resolution,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 3, pp. 1192–1204, 2022.
[12] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using Swin Transformer,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. Workshops (ICCVW), 2021, pp. 1833–1844.
[13] R. Timofte et al., “NTIRE 2017 challenge on single image super-resolution: Methods and results,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2017, pp. 1110–1121.
[14] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004.
[15] M. Bevilacqua, A. Roumy, C. Guillemot, and M. L. Alberi-Morel, “Low-complexity single-image super-resolution based on nonnegative neighbor embedding,” in Proc. Brit. Mach. Vis. Conf. (BMVC), 2012, pp. 1–10.
[16] R. Zeyde, M. Elad, and M. Protter, “On single image scale-up using sparse-representations,” in Proc. Int. Conf. Curves Surf., 2010, pp. 711–730.
[17] D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Proc. 8th IEEE Int. Conf. Comput. Vis. (ICCV), 2001, pp. 416–423.
[18] J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self-exemplars,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2015, pp. 5197–5206.
[19] Y. Matsui et al., “Sketch-based manga retrieval using Manga109 dataset,” Multimedia Tools Appl., vol. 76, no. 20, pp. 21811–21838, 2017.
[20] B. Lim, S. Son, H. Kim, S. Nah, and K. M. Lee, “Enhanced deep residual networks for single image super-resolution,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2017, pp. 1132–1140.
[21] Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, and Y. Fu, “Image super-resolution using very deep residual channel attention networks,” in Proc. Eur. Conf. Comput. Vis. (ECCV), 2018, pp. 286–301.