Registered datasets
Files already in your object store, read in place.
A registered dataset points at files that already sit in one of your tenant's object stores, such as a Delta table or a folder of Parquet exports that another tool writes. Beetl reads the files where they are and copies nothing. It never writes to, moves or deletes them.
Any user can register a dataset. The files must be in an object store the tenant already has, usually a bucket a TenantAdmin added (see Use your own bucket).
Register a dataset
On Home, right-click the canvas and choose Register data set, then pick the object store under "Where is the data?". Clicking an object store's node on Home opens the same dialog on that store.
Browse to the data. Click a file to register that one file, or open a folder and choose Use this
directory to register everything in it. For a Delta table, use the folder that contains
_delta_log. The browser shows "Delta table detected" when you are in it.
Check the name. Beetl suggests the file name without its extension, or the folder name. You write this name in SQL, and it follows the usual naming rules.
Choose Register data set. Beetl reads the location, detects the format and schema, and adds an Initialized dataset to Home, ready to query. If Beetl cannot read the location, no dataset is created and the dialog shows the reason.
Object keys that contain %, ?, #, spaces or control characters are greyed out in the browser
and cannot be registered. Rename them in your bucket first.
Supported formats
| Format | Point at | Beetl reads |
|---|---|---|
| Delta Lake | The table folder | The latest version of the table, with its partition columns |
| Parquet | One .parquet file, or a folder of them | Every Parquet file at that location |
| CSV | One .csv file, or a folder of them | Comma-separated files with a header row |
| JSON | One .json file, or a folder of them | Newline-delimited JSON, one object per line |
Beetl detects the format for you. A folder that contains _delta_log is read as Delta. Any other
location is tried as Parquet, then CSV, then JSON, and the first format that opens is used, so keep
one format per folder. A folder includes the files in its subfolders. Extensions must be lowercase,
and a JSON file whose content is one top-level array is refused.
Partition columns are recorded for Delta tables and marked in the schema on Home. CSV, Parquet and JSON datasets have none.
Where you can register from
A location cannot overlap anywhere Beetl writes its own data in that store: the folders of datasets Beetl writes, File Set uploads and saved query results. A path equal to, inside or above one of them is refused. Each location can be registered once per store, so registering the same path again fails until the first dataset is deleted.
Using a registered dataset
A registered dataset is a Bronze dataset you can query, preview, chart, mention in chat and use as a pipeline source. Beetl never writes to it:
- It cannot be a pipeline destination, and no sync, webhook or File Set can write to it.
- It never triggers a pipeline run. A pipeline that reads it runs when you press Run, when you deploy it, or when one of its other sources triggers it.
- In a query result built from it, drill-down to the records behind a cell is not available.
When the files change
Beetl reads the location every time a query, preview or pipeline run uses the dataset. New files in a registered folder, and new versions of a Delta table, appear in the next read without any action from you.
Beetl does not watch the location. A change starts no pipeline run, and the schema on Home stays the one Beetl read at registration. If the files are moved or deleted, reads of the dataset fail until they are back.
Renaming and deleting
Any user can delete a registered dataset, which leaves your files untouched. To change its name, contact Beetl; see Finding, renaming and deleting.
Related
Last updated on