Keep one copy of your data that every team can use.

Ask most companies where their customer data lives and you’ll get a list: the CRM, the warehouse, a Spark platform, two BI tools and a few spreadsheets. Each copy is a little different, and somebody has to decide which one to believe.

We wanted one answer. Your data lands once, as open Apache Iceberg tables that any engine can read, so you’re never locked in. One set of rules decides who sees what, whether they ask from a Data App, Tableau or a notebook.

Every tool reads the same tables.

In the usual stack, every tool keeps its own copy of the same table: the ingestion tool, the warehouse, the Spark platform, the BI extract, the app database. Every copy can drift, and every copy needs its own permissions.

Databasin keeps one. Pipelines, Data Apps, Databasin One and your BI tools all read the same tables, so the number in the board pack is the number on the dashboard. We’ve been in plenty of meetings where two people brought two different answers to the same question. One copy is how we fix that.

Copies of the same customer table In the usual stack the same table is copied about five times, once per tool; in Databasin there is one copy. Illustrative. 0 1 2 3 4 5 The usual stack The usual stack: 5 copies 5 copies Databasin Databasin: 1 copy 1 copy
Illustrative: copies of one customer table across five tools that each keep their own.

Compute scales to zero when nobody is using it.

Clusters scale down to nothing when nobody is using them and wake again when the work comes back. Overnight and on weekends, apart from scheduled jobs, they’re off, so you aren’t paying for an idle warehouse. The trade-off: the first query after a long quiet spell waits while the cluster starts up.

Minutes the cluster ran, per three-hour block, over one week Weekdays the cluster runs mostly between 9 AM and 6 PM, plus a short nightly pipeline; outside that it is asleep, apart from the short scheduled sync each night. Illustrative. 60 min 120 min 180 min 0 Mon Mon 03:00 to 06:00: running 15 of 180 minutes Mon 06:00 to 09:00: running 30 of 180 minutes Mon 09:00 to 12:00: running 150 of 180 minutes Mon 12:00 to 15:00: running 160 of 180 minutes Mon 15:00 to 18:00: running 120 of 180 minutes Mon 18:00 to 21:00: running 20 of 180 minutes Tue Tue 03:00 to 06:00: running 15 of 180 minutes Tue 06:00 to 09:00: running 30 of 180 minutes Tue 09:00 to 12:00: running 150 of 180 minutes Tue 12:00 to 15:00: running 160 of 180 minutes Tue 15:00 to 18:00: running 120 of 180 minutes Tue 18:00 to 21:00: running 20 of 180 minutes Wed Wed 03:00 to 06:00: running 15 of 180 minutes Wed 06:00 to 09:00: running 30 of 180 minutes Wed 09:00 to 12:00: running 150 of 180 minutes Wed 12:00 to 15:00: running 160 of 180 minutes Wed 15:00 to 18:00: running 120 of 180 minutes Wed 18:00 to 21:00: running 20 of 180 minutes Thu Thu 03:00 to 06:00: running 15 of 180 minutes Thu 06:00 to 09:00: running 30 of 180 minutes Thu 09:00 to 12:00: running 150 of 180 minutes Thu 12:00 to 15:00: running 160 of 180 minutes Thu 15:00 to 18:00: running 120 of 180 minutes Thu 18:00 to 21:00: running 20 of 180 minutes Fri Fri 03:00 to 06:00: running 15 of 180 minutes Fri 06:00 to 09:00: running 30 of 180 minutes Fri 09:00 to 12:00: running 150 of 180 minutes Fri 12:00 to 15:00: running 160 of 180 minutes Fri 15:00 to 18:00: running 120 of 180 minutes Fri 18:00 to 21:00: running 20 of 180 minutes Sat Sat 03:00 to 06:00: running 15 of 180 minutes Sun Sun 03:00 to 06:00: running 15 of 180 minutes Asleep, apart from the nightly sync.
Illustrative: minutes a cluster runs in each three-hour block of a typical week. Hover a bar for the block.
See the numbers
BlockWeekday minutesWeekend minutes
00–0300
03–061515
06–09300
09–121500
12–151600
15–181200
18–21200
21–2400

Three questions we hear in every first call

The IT lead Where does our data actually live?

With us: we run it, and your data is stored as open Iceberg and Delta tables, a format any engine can read, so you’re never locked in to us. If your rules say it has to stay inside your own Azure, we can install Databasin there instead.

Whoever pays the bill What are we paying for at 2am?

Nothing, if nothing is running. Compute is billed by the minute while it runs, with no seats and no annual commit.

The analyst Which table am I supposed to trust?

Raw data lands untouched, then gets cleaned, then becomes the business-ready tables people actually query. Each step is kept, so you can always see where a number came from.

Your metric definitions, applied everywhere.

Your data team writes the definitions once, in a Data Exchange: each metric, what it means, and how it’s calculated. Databasin One and Data Apps follow them when they answer, so two people asking the same question get the same number. Your data team owns the definitions, because they’re the ones people ask when a number looks wrong.

A Data Exchange listing approved metric definitions such as Abnormal Lab Results Count, Active Patient Count and Average Days to Payment. Metric definitions in a Data Exchange, each one approved and checked.

SQL and notebooks for your analysts

Analysts get a proper SQL editor and Python notebooks on the same tables everyone else is asking questions of. Write Spark SQL or PySpark against Databasin Sail, or plain SQL on Trino, and see the results right under the code.

The Databasin SQL editor: catalogs on the left, tabbed and versioned queries in the middle, joining project and financial tables in the silver layer.
The SQL editor: your catalogs on the left, tabbed queries with drafts and versions in the middle.
A Databasin notebook running a PySpark cell on Databasin Sail: a group-by over a lineitem table returning 2,526 rows in 4.13 seconds, shown as a table with filter, summary, CSV, JSON and chart options.
A notebook: a PySpark cell on Databasin Sail, with the results, a summary and a chart one click away. Mix Python and SQL cells in the same notebook.

For the technical readerHow the lakehouse is builtDatabasin SailSecurity and deployment

Connect one source and see where it lands.

$50 credit · No card · Your data in an open format, never locked in