Full Report
The volume of valuable data that organizations have to manage and analyze is growing at an incredible rate. This data is increasingly distributed across many locations, including data warehouses, data lakes, and NoSQL stores. As an organization’s data gets more complex and proliferates across disparate data environments, silos emerge, creating increased risk and cost, especially when that data needs to be moved. Our customers have made it clear; they need help. That’s why today, we’re excited to announce BigLake, a storage engine that allows you to unify data warehouses and lakes. BigLake gives teams the power to analyze data without worrying about the underlying storage format or system, and eliminates the need to duplicate or move data, reducing cost and inefficiencies. With BigLake, users gain fine-grained access controls, along with performance acceleration across BigQuery and multicloud data lakes on AWS and Azure. BigLake also makes that data uniformly accessible across Google Cloud and open source engines with consistent security. BigLake extends a decade of innovations with BigQuery to data lakes on multicloud storage, with open formats to ensure a unified, flexible, and cost-effective lakehouse architecture. BigLake architecture BigLake enables you to:Extend BigQuery to multicloud data lakes and open formats such as Parquet and ORC with fine-grained security controls, without needing to set up new infrastructure.Keep a single copy of data and enforce consistent access controls across analytics engines of your choice, including Google Cloud and open-source technologies such as Spark, Presto, Trino, and Tensorflow.Achieve unified governance and management at scale through seamless integration with Dataplex.Bol.com, an early customer using BigLake, has been accelerating analytical outcomes while keeping their costs low:“As a rapidly growing e-commerce company, we have seen rapid growth in data. BigLake allows us to unlock the value of data lakes by enabling access control on our views while providing a unified interface to our users and keeping data storage costs low. This in turn allows quicker analysis on our datasets by our users.”—Martin Cekodhima, Software Engineer, Bol.comExtend BigQuery to unify data warehouses and lakes with governance across multicloud environmentsBy creating BigLake tables, BigQuery customers can extend their workloads to data lakes built on Google Cloud Storage (GCS), Amazon S3, and Azure data lake storage Gen 2. BigLake tables are created using a cloud resource connection, which is a service identity wrapper that enables governance capabilities. This allows administrators to manage access control for these tables similar to BigQuery tables, and removes the need to provide object store access to end users. Data administrators can configure security at the table, row or column level on BigLake tables using policy tags. For BigLake tables defined over Google Cloud Storage, fine grained security is consistently enforced across Google Cloud and supported open-source engines using BigLake connectors. For BigLake tables defined on Amazon S3 and Azure data lake storage Gen 2, BigQuery Omni enables governed multicloud analytics by enforcing security controls. This enables you to manage a single copy of data that spans BigQuery and data lakes, and creates interoperability between data warehousing, data lake, and data science use cases. Open interface to work consistently across analytic runtimes spanning Google Cloud technologies and open source engines Customers running open source engines like Spark, Presto, Trino, and Tensorflow through Dataproc or self managed deployments can now enable fine-grained access control over data lakes, and accelerate the performance of their queries. This helps you build secure and governed data lakes, and eliminate the need to create multiple views to serve different user groups. This can be done by creating BigLake tables from a supported query engine like Spark DDL, and using Dataplex to configure access policies. These access policies are then enforced consistently across the query engines that access this data - greatly simplifying access control management. Achieve unified governance & management at scale through seamless integration with DataplexBigLake integrates with Dataplex to provide management-at-scale capabilities. Customers can logically organize data from BigQuery and GCS into lakes and zones that map to their data domains, and can centrally manage policies for governing that data. These policies are then uniformly enforced by Google Cloud and OSS query engines. Dataplex also makes management easier by automatically scanning Google Cloud storage to register BigLake table definitions in BigQuery, and makes them available via Dataproc Metastore. This helps end users discover these BigLake tables for exploration and querying using both OSS applications and BigQuery. Taken together, these capabilities enable you to run multiple analytic runtimes over data spanning lakes and warehouses in a governed manner. This breaks down data silos and significantly reduces the infrastructure management, helping you to advance your analytics stack and unlock new use cases.What’s next?If you would like to learn more about BigLake, please visit our website. Alternatively, get started with BigLake today by using this quickstart guide, or contact the Google Cloud sales team. Related Article Limitless Data. All Workloads. For Everyone Read about the newest innovations in data cloud announced at Google Cloud’s Data Cloud Summit. Read Article
Analysis Summary
# Industry News: Google Cloud Launches BigLake to Unify Multicloud Data Governance
## Summary
Google Cloud has announced the launch of **BigLake**, a new storage engine designed to unify data warehouses and data lakes across multicloud environments. The solution addresses the growing challenge of data silos by allowing organizations to analyze data across Google Cloud, AWS, and Azure without moving or duplicating it, while maintaining consistent security and governance.
## Key Details
- **Date:** April 6, 2022
- **Companies Involved:** Google Cloud (Primary), AWS, Azure (Interoperability partners), Bol.com (Early adopter)
- **Category:** Product Launch / Data Infrastructure & Security
## The Story
As data volumes explode and distribute across disparate environments (warehouses, lakes, NoSQL stores), organizations face increased operational costs and security risks associated with data movement. BigLake is Google’s answer to the "Lakehouse" architecture.
It functions as a storage engine that extends Google BigQuery’s capabilities to multicloud storage (Amazon S3, Azure Data Lake Storage Gen 2) and open formats (Parquet, ORC). By using a "service identity wrapper," BigLake enables fine-grained access control at the table, row, and column levels across various analytics engines—including open-source tools like Spark, Presto, and TensorFlow—without requiring the movement of the underlying data.
## Business Impact
### For the Companies Involved
- **Google Cloud:** Strengthens its "Data Cloud" value proposition, positioning itself as the central governance layer for organizations that are not exclusively on Google Cloud.
- **Bol.com:** Reported accelerated analytical outcomes and lower storage costs by eliminating the need for multiple data copies to achieve access control.
### For Competitors
- **Snowflake & Databricks:** BigLake directly competes with Snowflake’s Unistore/external tables and Databricks’ Lakehouse platform by offering similar cross-cloud, cross-format functionality.
- **AWS & Azure:** While Google is supporting their storage services, it is also attempting to "capture" the compute and governance layer for data residing on those platforms.
### For Customers
- **Reduced Overhead:** Eliminates the "data tax" of moving data between clouds for analysis.
- **Operational Efficiency:** Teams can use their preferred engines (Spark, Trino, etc.) under a single security umbrella.
- **Improved Compliance:** Centralized policy management makes auditing and data privacy enforcement significantly easier.
### For the Market
- This signals a shift toward **decoupled storage and compute** where governance is the primary value driver. It reinforces the industry trend toward "multicloud-native" architectures rather than cloud-locked silos.
## Technical Implications
BigLake utilizes **cloud resource connections** to act as a bridge between the query engine and the object store. It integrates with **Google Dataplex** for automated discovery and metadata management. For security, it leverages **Policy Tags**, which allow administrators to enforce security at a granular level that was previously difficult to achieve in unstructured data lakes.
## Strategic Analysis
- **Market Positioning:** Google is positioning itself as the most "open" and "interoperable" of the big three cloud providers by supporting open-source engines and competitor storage.
- **Competitive Advantage:** The ability to apply BigQuery-level security to files in Amazon S3 or Azure Gen 2 provides a significant governance advantage over native tools.
- **Challenges:** Performance overhead across clouds (egress/latency) and the complexity of managing cross-cloud service identities may pose adoption hurdles for smaller firms.
## Industry Reactions
- **Analyst Opinions:** Generally viewed as a necessary evolution for Google to maintain BigQuery’s dominance in a multicloud world.
- **Expert Commentary:** Observers note that this lowers the barrier for "Best-of-Breed" stack building, as companies no longer have to commit to one cloud provider's full analytics suite.
## Future Outlook
- Expect Google to deepen integration with third-party security vendors to further automate policy tagging.
- The success of BigLake will likely force AWS and Azure to release more robust, cross-cloud governance tools for their own data lake services.
## For Security Professionals
BigLake is highly relevant to CISO and Data Protection offices because it addresses the **"Shadow Data"** problem. By enforcing row- and column-level security across different clouds and engines from a single point (Dataplex), security teams can significantly reduce the risk of unauthorized data exposure in lakes. It allows for the implementation of **Least Privilege** access to data files without needing to manage complex IAM policies across three different cloud providers individually.