# Databricks

![Databricks_logo.png|400](https://assets.relyanceuat.xyz/images/docs/5779042266253/5778993664397_white.png)

Databricks is a cloud-based data engineering tool used for processing and transforming massive quantities of data and exploring the data through machine learning models. Developed by the creators of Apache Spark, it provides a unified, open platform for all your data.

You can connect Databricks to Relyance AI in two ways. Pick one under **Authentication Method** on the connection wizard's **Authentication** step:

- **OAuth (service principal)**, recommended. A Databricks service principal authenticates with OAuth machine-to-machine credentials. It gets only the privileges you grant it, and nothing depends on an individual user's personal token.
- **Custom**, with a **Databricks Instance URL** and a **Personal Access Token**. This is the original method and keeps working for existing connections.

## Connect with a service principal (recommended)

### In Databricks

A workspace admin does these steps.

1. **Note the workspace URL.** This is the **Databricks Instance URL**, for example `https://adb-1234567890123456.7.azuredatabricks.net`.
2. **Create a service principal.** In the workspace, open **Settings** → **Identity and access** → **Service principals**, click **Add service principal**, and give it a name such as `relyance-scanner`.
3. **Generate an OAuth secret.** Open the service principal, go to its **Secrets** tab, and click **Generate secret**. Copy the **Client ID** and the **Secret** right away; the secret is shown only once. Databricks OAuth secrets expire, so note the expiry date.
4. **Grant access to the data you want scanned.** In **Catalog Explorer**, grant the service principal `USE CATALOG` on each catalog, `USE SCHEMA` on each schema, and `SELECT` on the tables to scan. Granting at the catalog level covers everything below it.
5. **Make it a workspace admin** (add it to the workspace `admins` group). Relyance needs this to list notebooks and the workspace's users, groups and service principals.
6. **Optional: give it a SQL warehouse to sample rows.** Grant the service principal **CAN USE** on a SQL warehouse and copy the warehouse's **ID** from its **Connection details**. Without a warehouse, Relyance scans metadata only (catalogs, schemas, tables, columns and grants) and reads no row data.

### In the Relyance AI application

1. Log in to your Relyance AI account.
2. Open **Settings** (bottom left) and select **Integrations**.
3. Find the **Databricks** integration card and click it, then click **Add Connection**.
4. Follow the wizard. On the **Authentication** step, choose **OAuth (service principal)** and fill in:
    - **Databricks Instance URL**: the workspace URL from step 1.
    - **Client ID** and **Client Secret**: from step 3.
    - **SQL Warehouse ID** (optional): from step 6. Letters and digits only, for example `1234567890abcdef`. Leave it blank for a metadata-only scan.
    - **Config Options** (optional): which catalogs, schemas and objects to scan. See [Limit what is scanned](#limit-what-is-scanned) below. Leave the default to scan everything the service principal can see.
5. Click **Authenticate** and finish the wizard. Relyance checks each step separately and tells you which one failed: getting a token, listing catalogs, and reaching the SQL warehouse if you set one.

## Limit what is scanned

**Config Options** takes an allow list and a deny list of Unity Catalog names:

```json
{
  "allow_list": [
    "prod",
    "prod.analytics.*"
  ],
  "deny_list": [
    "prod.hr"
  ]
}
```

- **Patterns** have one to three dot-separated parts: `catalog`, `catalog.schema` or `catalog.schema.object`. Each part may use `*` (any characters) and `?` (one character). Matching is case-insensitive, and missing trailing parts match everything, so `prod` means every schema and object in the `prod` catalog.
- **An object is scanned** when the allow list is empty or any allow pattern matches it, **and** no deny pattern matches it. Deny wins.
- **Scope:** the lists apply to catalogs, schemas, tables, columns, volumes, functions, registered models and their grants. Notebooks, jobs, pipelines, workspace identities, model serving endpoints, Genie spaces, apps and experiments are not filtered.
- **Invalid lists are rejected.** Unknown keys such as a misspelled `denylist`, a value that isn't a list of strings, a pattern containing whitespace, an empty part, or more than three parts all fail validation. Relyance never falls back to scanning everything.
- **Excluded objects are removed from inventory.** Objects that fall outside the lists stop appearing in your inventory after the next scan. Widen the lists again and they come back on the following scan.

Examples:

| Config Options | Scans |
| --- | --- |
| `{"allow_list": [], "deny_list": []}` (default) | everything the service principal can see |
| `{"allow_list": ["prod"]}` | only the `prod` catalog |
| `{"allow_list": ["*.sales"]}` | the `sales` schema in every catalog |
| `{"deny_list": ["dev_*", "prod.scratch"]}` | everything except catalogs starting with `dev_` and the `prod.scratch` schema |

## Connect with a personal access token

**Token based authentication** must be enabled by the administrator.

**In Databricks:**

1. The **Databricks instance URL** will be the custom Databricks URL for your organization.
2. Click **Settings** in the lower left corner of your Databricks workspace.
3. Click **User Settings**.
4. Go to the **Access Tokens** tab.
5. Click the **Generate New Token** button.
6. Click the **Generate** button and copy the generated token.

**In the Relyance AI application:**

1. Login to your Relyance account.
2. Navigate to the **Settings** Menu in the bottom left-hand side.
3. Select **Integrations**.
4. Find the **Databricks** integration card and click on it.
5. In the Databricks integration page, click **Add Connection**.
6. Follow the wizard, adjusting any defaults as needed.
7. Choose **Custom** for the Authentication Method.
8. Paste the **Databricks Instance URL** and **Personal Access Token** into their respective fields. You can also specify the **Clusters** and **SQL Warehouses** to scan. **For Asset Discovery its important and necessary to provide the warehouse or cluster name.**

![Screenshot](https://assets.relyanceuat.xyz/images/docs/5779042266253/42979366201613.png)

![Screenshot](https://assets.relyanceuat.xyz/images/docs/5779042266253/42979366207629.png)

![Screenshot](https://assets.relyanceuat.xyz/images/docs/5779042266253/42979366208397.png)

9. Click **Authenticate** and finish the wizard.
10. At this point, you should see the following result on the Vendor Integrations page:
11. Congratulations, you are now connected to **Databricks**.

![Databricks-2.png](https://assets.relyanceuat.xyz/images/docs/5779042266253/5780632758285.png)

<!-- failure-modes:begin (generated from the integration catalog; do not hand-edit) -->

## If the connection reports Connected but returns nothing

These are the ways this integration comes back empty without reporting an error. Generated from the integration catalog, so it tracks what the connection actually asks for.

1. **A credential rotated at the vendor is not picked up here.** **Personal Access Token** and **Client Secret** are stored when you save the connection, so regenerating the value at the vendor breaks the next scan until it is re-pasted here. Recording the expiry on the connection means Relyance warns you before it lapses.
2. **Check the address fields before suspecting the credentials.** **Databricks Instance URL** identifies which tenant, region or host to talk to. A wrong value there fails authentication and looks exactly like a bad secret.

<!-- failure-modes:end -->

<!-- auth-methods:begin (generated from the integration catalog; do not hand-edit) -->

## Authentication methods and fields

Pick one of these under **Authentication Method** on the connection wizard's **Authentication** step. This table is generated from the integration catalog, so it always matches what the form actually asks for.

| Method | Required | Optional |
| --- | --- | --- |
| **Custom** | `Databricks Instance URL`, `Personal Access Token` (secret), `Cluster Options`, `SQL Warehouse Options` | — |
| **OAuth (service principal)** | `Databricks Instance URL`, `Client ID`, `Client Secret` (secret) | `SQL Warehouse ID`, `Config Options` |

<!-- auth-methods:end -->

<!-- terraform-examples:begin (generated from the integration catalog; do not hand-edit) -->

## Manage this integration with Terraform

Connections for this integration can be managed as code with the [Relyance Terraform provider](https://registry.terraform.io/providers/Relyance/relyance/latest). Non-secret fields go in `auth.params`; secret fields go in `auth.secrets_wo`, which is write-only — never stored in Terraform state. Rotate secrets by bumping `auth.secrets_wo_version`.

### API token

```hcl
resource "relyance_integration_connection" "databricks_0" {
  vendor = "databricks"
  name   = "<your connection name>"

  auth = {
    method = "api-token"
    params = {
      instance_url = "<instance_url>"
      clusters = jsonencode({
        cluster_names = [
          ""
        ]
      })
      sql_warehouses = jsonencode({
        sql_warehouse_names = [
          ""
        ]
      })
      data_storage_location = "us"
    }
    # Secret fields are write-only: sent to Relyance, never stored in state.
    secrets_wo = {
      api_token = var.databricks_api_token
    }
    secrets_wo_version = 1
  }

  scans = { "data-inspection" = { enabled = true } }
}
```

### Client ID & secret

```hcl
resource "relyance_integration_connection" "databricks_1" {
  vendor = "databricks"
  name   = "<your connection name>"

  auth = {
    method = "client-credentials"
    params = {
      instance_url = "<instance_url>"
      client_id = "<client_id>"
      data_storage_location = "us"
    }
    # Secret fields are write-only: sent to Relyance, never stored in state.
    secrets_wo = {
      client_secret = var.databricks_client_secret
    }
    secrets_wo_version = 1
  }

  scans = { "data-inspection" = { enabled = true } }
}
```

<!-- terraform-examples:end -->
