Google Cloud expands borderless Lakehouse across clouds
Thu, 30th Jul 2026 (Today)
Google Cloud has introduced new features for its borderless Lakehouse, extending data access across on-premises systems, rival clouds and software applications.
The changes centre on Google's use of the Apache Iceberg format and are intended to let customers query data where it already sits, rather than moving it into a single repository. The model supports access to operational systems across clouds, as well as software platforms used for finance, human resources and customer management.
A key part of the launch is catalog federation, now in preview for AWS Glue, Databricks Unity and Snowflake Horizon. Through those links, customers can discover and query data held in other environments through BigQuery, Google's managed Spark service and other Iceberg-compatible engines.
That means a company using multiple cloud providers can read the same data in place instead of creating duplicate copies in separate analytics systems. The arrangement also allows data to be written back across environments after processing or analysis.
Google is also extending the approach to software applications including SAP, Salesforce and Workday. According to Google, BigQuery can query live application data in those platforms directly, while the applications can also use BigQuery AI functions on data that remains in place.
Cross-cloud links
Another part of the announcement is the use of Cross-Cloud Interconnects between Google Cloud and other providers. These private links are intended to reduce the cost and delay of moving large datasets between clouds for analytics and machine learning workloads.
For customers accessing data from AWS, the borderless Lakehouse supports zero variable egress costs, though users still pay an hourly fee for interconnection services. Google is offering the service through a subscription model covering links from 1G to 100G.
The platform also uses cross-cloud caching to store fragments of remote data temporarily inside Google Cloud. This is meant to avoid repeated transfers when customers run follow-up business intelligence or ad hoc queries against the same source data.
The broader pitch reflects a growing effort by cloud providers to position analytics systems as foundations for AI tools and software agents. In that model, companies want language models and automated systems to work across fragmented data estates without relying on long-running extract, transform and load pipelines.
The borderless Lakehouse is intended to support Gemini Enterprise and conversational agents that can analyse data regardless of physical location. Google is pairing that effort with its Data Agent Kit and Conversational Analytics API, aimed at developers building custom agents that can operate across datasets and return results in natural language.
Governance layer
To address concerns about context and trust in AI outputs, Google is linking the platform to its Knowledge Catalog service. The catalog can synchronise metadata from AWS Glue, Databricks Unity Catalog and Snowflake Horizon, then index and translate technical schemas into business terminology with lineage information.
That governance layer is intended to give users table-level access control and support credential vending across platforms. It should also help agents understand what data they are permitted to use and provide a common semantic layer across a company's information estate.
The launch also broadens the borderless Lakehouse concept into Google's database portfolio. Spanner Omni, which runs outside Google Cloud, and Lakehouse Federation for AlloyDB are part of the expansion, allowing transactional systems and analytics environments to connect more directly.
Google argues that the economics of the model rest on less data movement, fewer bespoke pipelines and tighter control over AI-related computing costs. It says Knowledge Catalog can reduce the amount of context sent to models, while BigQuery AI includes controls for estimating and limiting token use before queries run.
According to Google, customers using BigQuery's cost-optimised AI functions have seen a 230x reduction in token consumption. The figure is part of a broader effort by cloud vendors to show that AI analytics can be managed as a cost discipline as well as a technical one.
The announcement underlines how competition in enterprise data platforms is shifting from storage and processing speed towards interoperability, governance and AI access. Rather than persuading customers to centralise everything in one cloud, providers are increasingly trying to make their software useful in environments where data is spread across several clouds, databases and applications.
Google says the borderless Lakehouse now offers secure, bi-directional access to data across external catalogs and software platforms, while allowing analytics and AI tools to run against data in place.