arXiv · 2601.17546
Push Down Optimization for Distributed Multi Cloud Data Integration
Abstract
Enterprises increasingly adopt multi cloud architectures to take advantage of diverse database engines, regional availability, and cost models. In these environments, ETL pipelines must process large, distributed datasets while minimizing latency and transfer cost. Push down optimization, which executes transformation logic within database engines rather than within the ETL tool, has proven highly effective in single cloud systems. However, when applied across multiple clouds, it faces challenges related to data movement, heterogeneous SQL engines, orchestration complexity, and fragmented security controls. This paper examines the feasibility of push down optimization in multi cloud ETL pipelines and analyzes its benefits and limitations. It evaluates localized push down, hybrid models, and data federation techniques that reduce cross cloud traffic while improving performance. A case study across Redshift and BigQuery demonstrates measurable gains, including lower end to end runtime, reduced transfer volume, and improved cost efficiency. The study highlights practical strategies that organizations can adopt to improve ETL scalability and reliability in distributed cloud environments.
Explore related subjects
Keep this discovery
Ravi Kiran Kodali, Vinoth Punniyamoorthy, Akash Kumar Agarwal, Bikesh Kumar, Balakrishna Pothineni, Aswathnarayan Muthukrishnan Kirubakaran, Sumit Saha, Nachiappan Chockalingam. 2026-01-24. Push Down Optimization for Distributed Multi Cloud Data Integration. https://doi.org/10.5120/ijca2026926214
Cite the original work for its findings. Save a collection to share your selection of sources.