Hash Keys vs Identity Columns: Choosing Surrogate Keys in the Lakehouse
Deterministic hash surrogate keys or IDENTITY columns for your dimensions? A practitioner's pros and cons across load order, idempotency, join cost, and reprocessing.
Delta internals, secret management, and identity on Databricks.
3 articles
Deterministic hash surrogate keys or IDENTITY columns for your dimensions? A practitioner's pros and cons across load order, idempotency, join cost, and reprocessing.
When defining identity columns in Databricks tables, two common options for automatic identity value generation are GENERATED BY DEFAULT and GENERATED ALWAYS. This article explains the differences and when to use each.
Databricks can connect to various sources for data ingestion. This article describes how to manage Azure Key Vault-backed secret scopes in Databricks using the GUI.