How it works

From raw source
to cited answer.

One platform, five modules, no copies between systems. Click any module and go exactly as deep as you want — every design decision is documented.

01

Connect: live sources & pipelinesQuery it now, or sync it for keeps — your call per source

Live sources · new
Query it now. No sync.

Connect a supported source and it appears as live tables the moment credentials pass — SELECT * FROM live.hubspot.contacts runs directly against the API. First insight in minutes, zero pipeline, and your databases federate the same way. Built for exploration and quick answers.

Pipelines · sync
Sync it for keeps.

One button builds the pipeline: schema mapped, schedule set, changes tracked. Data lands in your own lakehouse — full history, fast queries at any volume, and pre-built gold views for the 30 native sources. Built for speed, scale, and everything downstream.

Both lanes run on the same connector: schema-aware, probe-validated, built to understand the source's actual data model — the Epic connector knows Chronicles, Clarity, and Caboodle as distinct environments; the Workday connector resolves business objects and effective-date logic before your data ever lands.

30 native, zero-config sources
Epic, Workday, Salesforce, Sage Intacct, HubSpot, Marketing Cloud and more — certified routes, semantic model included.
Everything else still connects
Generic REST/SOAP with a no-code API builder, plus JDBC for databases. Any HTTP endpoint becomes a source — no engineering ticket.
Schema drift stops at bronze
Upstream changes are versioned and absorbed at the raw layer — not cascaded into your reports at 2am. Pipelines adapt instead of failing.
Incremental by default
Watermarks, change-data-capture, retries, and full run history on every pipeline. Your OLTP systems never get polled for analytics.
The live lane — pick a source, make it live, query it immediately (real footage)
The sync lane — pipeline runs with live logs, including AI document embedding (real footage)
Databasin connector catalog — 75+ pre-built source connectors in one library
The connector catalog — 75+ sources, one library
Databasin pipelines dashboard showing run statuses and source-to-target wiring
Pipelines — run status, schedules, source-to-target wiring

Replaces CData · MuleSoft · Fivetran · custom API scripts

Live sources are rolling out across the native connectors now — the product shows you exactly which of your sources support live; everything supports sync. Browse all connectors →

02

Automate: the work engineSQL, dbt, notebooks, and AI agents — chained into stages

Connecting brings data in; Automations put it to work. A full orchestration engine built into the core, not a bolt-on: tasks chain into stages — tasks in the same stage run in parallel, stages run in order — and the whole board runs on a schedule or a trigger, with retries, per-task logs, and versioned snapshots of the automation itself.

SQL & dbt, first-class
Schedule plain-SQL transforms on any engine, or point a task at your dbt Core repo — profiles generated, per-model status reported back.
Notebooks on the lake
PySpark, Python, and Scala run directly against governed tables. Data scientists ship models, not tickets.
AI agents as tasks
Put a skill on a schedule — exec summary, data-quality check, anomaly explanation. Read-only tools, query budgets, full audit trail per run.
Semantic models in plain English
Describe the model you need; Databasin drafts it from your schemas, you refine it on a canvas, and gold views materialize from it.
Delivery built in
Results ship to email, Slack, or Teams, or write back to tables and external systems — reverse-ETL without another vendor.
Drives what you already own
Trigger Databricks jobs, hit APIs, fire webhooks — external systems become stages in the same workflow.
Adding a task — every task type on the menu, then an AI-agent task configured end-to-end (real footage)

Replaces Airflow / Dagster · dbt runners · Hightouch / Census · cron & scripts

03

Store & query: the open lakehouseApache Iceberg you own · Trino & Spark · zero copies

Synced data lives in Apache Iceberg — a vendor-neutral open format readable by any compatible engine, now and in the future. No exit tax, ever. Pipelines and automations move it through three governed tiers, and your gold layer is built from the questions you actually ask:

Bronze
Raw, as-ingested. Full history, nothing dropped.
Silver
Cleaned, typed, deduped, conformed.
Gold
Curated business views — joins, keys, definitions. What you query and what the AI answers from.

Two engines. One lake. Zero copies. Managed clusters bill per node-minute and cost nothing when stopped; access control runs per user, per catalog, per schema. And one Trino query can join it all — synced gold tables, live sources, and federated databases in a single statement.

Trino
Interactive SQL & federation
Low-latency, ad-hoc analytics against gold tables — plus live sources and federated Postgres, MySQL, SQL Server, Oracle, MariaDB, or Snowflake, joined in one statement. The engine behind dashboards and the SQL editor.
Apache Spark
Heavy processing & ML
Multi-terabyte transformations, nightly medallion refreshes, notebooks, and machine-learning pipelines — distributed compute against the same Iceberg tables, no infrastructure to manage.

Also in the toolbox: Apache Doris for sub-second, high-concurrency OLAP serving, DuckDB for in-process analytics — and BYO mode layers Databasin onto your existing Databricks, Snowflake, or Fabric instead of replacing it.

Interactive SQL on Trino — governed gold tables, answers in seconds (real footage)
Databasin SQL editor — interactive Trino query against governed gold-layer tables
The SQL editor — one surface for every engine
Databasin cluster management — Trino and Spark clusters, fully managed
Cluster management — per-minute billing, stopped clusters cost nothing

Replaces or overlays Snowflake · standalone Databricks · Azure Synapse · BigQuery

04

Ask: Databasin OneThe AI layer — answers with receipts, charts, documents, agents

The AI layer is architecturally constrained to your governed data — it can't hallucinate against raw tables because it never sees them. LLM-agnostic: GPT and Claude models via Azure OpenAI, or your internally approved model, running inside your security boundary. Governance and AI are built together, not bolted together.

Answers with receipts
Every answer ships the SQL it ran and the tables it touched. Trust it, or check it — either way, no arguments about whose number is right.
Documents cite the page
PDFs and documents are chunked and embedded inside the platform — nothing leaves. Ask a question, click the citation, see the highlighted source page.
Charts → dashboards → gallery
"Chart this" gives you options; tiles compose into live dashboards; published dashboards re-run as the viewer, with their permissions.
Documents & spreadsheets out
Executive PDFs with embedded charts, Excel exports, narrative summaries — generated from live data, deliverable on a schedule.
Agent skills, versioned
Skills are managed like code: versioned, shareable, admin-curated. Teams run them with one click; automations run them on a schedule.
Guardrails you can audit
Agents get read-only tools, query budgets, and row caps. Every run writes an audit trail — what it read, what it ran, what it produced.
A real session — question in, narrative + SQL + results + charts out (real footage)
Chart → dashboard → publish, without leaving the conversation (real footage)

Replaces Tableau · Power BI Premium · Looker · standalone AI API spend

05

Trust: governance & securityHIPAA-ready, encrypted, audited — trust by architecture

Databasin was co-created at Washington University School of Medicine, where the data was live, regulated PHI from day one. Security isn't a compliance checkbox layered on later — the architecture assumes the strictest posture and lets you relax it, not the reverse.

Encrypted & audited
Encryption in transit and at rest, full audit trails, SSO and role-based access control on every deployment.
Access control where the data lives
Permissions at the catalog, schema, and row level — enforced by the engines, not by the BI tool's honor system.
AI inside the boundary
The model queries governed gold data only, and runs inside your security perimeter. Sensitive records are never shipped to an external AI service.
Your tenant, if you need it
Self-install puts the identical platform inside your own Azure tenant — your network, your keys, your policies, no data egress.
Databasin governance — catalog and schema permissions, users, groups, and roles
Governance — catalog & schema permissions, users, groups, roles
Design principles

Every architectural decision has a reason.

Open

Apache Iceberg — a vendor-neutral open format. Your data is readable by any compatible engine, now and in the future. No exit tax.

→ No lock-in at the storage layer, ever.
Pluggable

Sources, engines, agent skills, and destinations all plug in — and swap out — without re-platforming. Full stack or single module, your call.

→ You choose the pieces; the platform doesn't choose for you.
Secure

Private install in your own Azure tenant when you need it. Your LLM, your endpoints, your governance rules. PHI never leaves your environment — by architecture, not by policy.

→ Sovereignty is a design requirement, not a checkbox.
Intelligent

AI woven into the architecture, not bolted on. The AI layer queries governed gold data only — governed data in, trusted answers out.

→ Intelligence requires governance. We build both together.
Run it your way

HIPAA-ready by default.
Your security posture, your call.

Every deployment is encrypted, audited, and access-controlled. The hosted cloud is fully HIPAA-ready — most teams start there in five minutes. For the strictest PHI and data-residency needs, run the identical platform inside your own Azure tenant.

The fast path.

Fully managed and HIPAA-ready from day one. Sign up, click a connector, and you're querying in minutes — $50 in credit, no card.

Try it now

Strictest posture.

Install from the Azure Marketplace — the whole platform inside your walls. Your storage, your keys, your network, no data egress. Unlimited seats.

Talk to us about self-install
HIPAA-ready
Encrypted in transit & at rest
Audit · row-level security
SSO · RBAC

Co-created at Washington University School of Medicine — built where the data was real, and regulated. Featured by Microsoft and Databricks at HIMSS '23, '24, and '25.

Five minutes · $50 credit · no card

See it work on your own data.

The whole path you just read — running on your sources in five minutes.

Try it now →
Five minutes, $50 in credit

See the whole path
on your own data.