API reference¶
duckrun connection API — supported methods¶
✅ 11 public methods · 62/62 tests passing
Introspected from the shipped classes — the exact public surface of
duckrun.connect(), signatures and all, not a hand-maintained list. The green suite (test_connection_api.py) vouches it works.conn.sql()also routes raw Delta DML — see the DML matrix on the Connection API page.
| Surface | Method | Parameters |
|---|---|---|
duckrun |
connect |
path, storage_options=None, schema=None, read_only=True, name=None, format='delta' |
DuckSession |
attach |
path, name=None, storage_options=None, schema=None, read_only=None, format='delta' |
DuckSession |
close |
(none) |
DuckSession |
convert_to_delta |
identifier, partition_schema=None |
DuckSession |
copy |
local_folder, remote_folder, file_extensions=None, overwrite=False |
DuckSession |
download |
remote_folder='', local_folder='./downloaded_files', file_extensions=None, overwrite=False |
DuckSession |
get_stats |
source=None, detailed=False |
DuckSession |
list_files |
remote_folder='', file_extensions=None |
DuckSession |
refresh |
quiet=False, catalog=None |
DuckSession |
register |
name, obj |
DuckSession |
sql |
query |
duckrun workspace API — Fabric artifact deploy¶
🗂️ 16 public methods · deploy · run · schedule
Introspected from the shipped classes — the exact public surface of
duckrun.workspace()and theWorkspacehandle it returns, signatures and all, not a hand-maintained list. It drives Microsoft Fabric (create lakehouses, deploy notebooks / semantic models / pipelines / variable libraries, run and schedule them) — see the Workspace (Fabric) page. Exercised by the manual deploy demo (tests/deploy_testing).
| Surface | Method | Parameters |
|---|---|---|
duckrun |
workspace |
workspace, token=None |
Workspace |
create_lakehouse |
name, schemas=True, folder=None |
Workspace |
create_warehouse |
name, folder=None |
Workspace |
deploy |
source, lakehouse=None, variables=None, name=None, overwrite=False, notebook=None, warehouse=None, folder=None, mode=None |
Workspace |
display_name |
property |
Workspace |
download |
folder='.', name=None, overwrite=False |
Workspace |
id |
accessor |
Workspace |
lakehouse_id |
name |
Workspace |
list_items |
kind=None |
Workspace |
list_lakehouses |
(none) |
Workspace |
name |
accessor |
Workspace |
run |
name |
Workspace |
run_python |
script, *, lakehouse=None, args=None, env=None, cores=None, pip=None, setup=None, entry=None, name=None, attempts=3, keep_notebook=False |
Workspace |
schedule |
name, every=None, daily=None, weekly=None, at=None, tz='UTC' |
Workspace |
sql_endpoint |
warehouse=None |
Workspace |
warehouse_id |
name |
Delta SQL extensions (not in DuckDB)¶
Everything you run through conn.sql() is standard DuckDB SQL — reads, CTEs, SHOW/DESCRIBE,
CREATE TEMP/CREATE VIEW, and the write DML (CREATE TABLE … AS, INSERT, UPDATE, DELETE,
MERGE, ALTER TABLE, DROP TABLE) are all written in ordinary DuckDB syntax; duckrun only routes
the writes to delta-rs so they land on the Delta table (see the
DML matrix).
The forms below are the only places where the SQL is not vanilla DuckDB — new verbs and clauses DuckDB has no syntax for (Spark/Delta spellings). A read-only session refuses the ones that write.
| Extension | What it does | Writes? |
|---|---|---|
CREATE [OR REPLACE] TABLE <t> SORTED BY AUTO AS <query> |
duckrun profiles the query and picks a run-length-friendly clustering key for you (a heuristic, not an optimizer). Only the AUTO keyword is the extension — SORTED BY (cols) and PARTITIONED BY (cols) are DuckDB's own CTAS syntax. See Automatic sorting. |
✅ |
INSERT INTO <t> REPLACE WHERE <pred> SELECT … |
delta_rs replaceWhere — atomically overwrite only the rows matching <pred> with the SELECT's rows, in one fenced commit (no torn delete-then-append window). <pred> is a CAST-free expression over the target's columns. |
✅ |
INSERT WITH SCHEMA EVOLUTION INTO <t> SELECT … |
append that widens the table with the source's new columns (existing rows → NULL) instead of dropping them — delta_rs schema_mode='merge'. |
✅ |
RESTORE TABLE <t> TO VERSION AS OF <n> RESTORE TABLE <t> TO TIMESTAMP AS OF '…' |
delta_rs restore — roll the table back to an earlier version/timestamp. It's a new commit on top of history, so the restore is itself revertible. |
✅ |
DESCRIBE DETAIL <table> |
the Delta table's location (storage path), partitionColumns, numFiles, sizeInBytes, version — read from the Delta log. Plain DESCRIBE <table> stays DuckDB's column view. |
— |
DESCRIBE HISTORY <table> |
one row per Delta commit (version, timestamp, operation, operationMetrics), newest first. |
— |
Not extensions — these are legit DuckDB verbs/syntax; duckrun only changes what they do against a Delta table, without adding any new spelling:
VACUUM <table>— DuckDB's ownVACUUMverb, repurposed to compact small files and vacuum files tombstoned past the retention window (compaction also runs automatically after every write; this is the manual button). BareVACUUMwith no table stays DuckDB's stats no-op.SORTED BY (cols)/PARTITIONED BY (cols)— DuckDB's native CTAS layout clauses, applied to the Delta write.- Time travel — DuckDB's own
delta_scan('<location>', version => N). Get<location>fromDESCRIBE DETAILand the versions fromDESCRIBE HISTORY.
Portability: everything above except the extensions table is plain DuckDB SQL — the same query runs on stock DuckDB. The full write-DML behaviour is pinned statement-by-statement in
tests/connection_api/test_connection_api.py.