High-fidelity scientific instruments and simulations are producing data at unprecedented volumes and rates, imposing substantial pressure on storage, transmission, and analysis. In order to alleviate the challenges brought to high-performance computing I/O and storage, compression techniques have been introduced in various scenarios, but understanding compression performance remains challenging due to the complex interactions among data characteristics and compressor-internal behaviors. In this paper, we analyze the internal behaviors of SZ, a representative prediction-based error-bounded lossy compressor, and identify representative compressor-related features for performance prediction. We then develop learning-based models to predict compression ratio and throughput, and further design a simplified prediction model using representative features. We compare the proposed method with existing sample-based and white-box prediction methods. The results show that compressor-internal features are important for performance prediction, but the best prediction method is scenario-dependent.



