Enter the duck#
This brings us to DuckDB– a whimsically named, in-process, analytical database. DuckDB is maintained by DuckLabs, a company with academic roots in the Centrum Wiskunde & Informatica, a research institute in Amsterdam and the birthplace of Python. You can run it just about anywhere by downloading a single file, without the irritating setup. Listen closely, and you can hear DuckDB fans frothing with joy: “it just works!”
DuckDB has benefited bigly from the rise of coding agents. From the AWS blog announcing their acquisition of DuckLabs:
“And what works for everyday queries also (unsurprisingly) works very well for agents because agents behave a lot like people when interacting with data. They poke. They experiment. They run exploratory analysis on small data sets before figuring out what they really want to do. DuckDB ends up being naturally optimized for AI agents to use.”
Point Claude Code at a huge CSV file and you’re likely to see installing duckdb in your terminal. It’s fast, powerful, and simple for agents to use without much user direction.
Human or clanker, DuckDB users have a lot in common. Here are a few of the most popular use cases.
Local analytics on files#
DuckDB shines for local analytics, using the processing power of your local machine to explore and analyze datasets. DuckDB can query files directly in a variety of formats without requiring the user to load the file into a database first, making the analytics lifecycle faster than setting up a database server, building an ingestion pipeline, attaching your client, and finally writing a query.
For example, say you have a file of events occurring in a web application:
The data itself can exist in a wide variety of tabular formats: comma-separated values (CSV), Parquet files, or good old-fashioned Excel. It could be physically on your local machine, served via an API, or hosted in cloud object storage like Amazon S3.
DuckDB doesn’t care. You can answer questions like “how many pageviews landed on the pricing page?” without moving data yourself–DuckDB handles it for you. This is a big ducking deal. You don’t need to run a database server and build an ingest pipeline just to start exploring a dataset.
Notice how in the query below, DuckDB reads directly from an S3 storage bucket. This is just one example of DuckDB’s simplicity: querying a remote dataset using the processing power already on your local machine.
SELECT COUNT(*) AS pricing_pageviews
FROM 's3://my-bucket/events.csv'
WHERE event_type = 'page_view'
AND page = '/pricing';
“Where do I actually run this SQL?” you ask. You can use DuckDB just about anywhere, from Jupyter notebooks to the command line interface, or even in a browser. Or, if you don’t want to think about installations or SQL syntax at all, just tell your coding agent to do it. Claude, Codex, and company are adept users of DuckDB.
In terms of scale, DuckDB can comfortably crunch through hundreds of gigabytes (Kim K’s Insta data, probably) on most modern MacBooks. Of course, a beefier machine helps, but it isn’t necessary. The DuckLabs team even completed some pretty gnarly benchmarks running DuckDB on a MacBook Neo–Apple’s entry-level laptop with a meager 8GB of memory. Thanks to DuckDB’s excellent larger-than-memory workload processing, it can tame files that are much larger than the memory on your local machine.
In-browser queries with WebAssembly#
Some web applications are slow, and some are fast. Historically, business intelligence tools are some of the slow ones.
Yes, older tools can be fast, given the right tuning. But, it depends(™). In many cases, the query engine still lives on the server–no matter how fast it is, your experience of speed will always be capped by requests and responses flying back and forth over the network. Even fast systems like Snowflake and Google BigQuery can only return results as fast as the network will allow.
The über-fast, interactive web apps you use nowadays (looking at you, Figma) often use WebAssembly. WebAssembly, or Wasm, lets you run code from fast languages like C++ or Rust directly in the browser, including DuckDB!
DuckDB-Wasm changes where query execution happens. Interactions like filtering and scrubbing run right on your machine instead of round-tripping to a server. These are real queries, run by DuckDB in the browser itself. This is only possible because DuckDB is an in-process analytical database—you cannot stuff Snowflake into the browser.
A common pattern is to run an initial query on the server-side database engine, persist the results as a table in DuckDB-Wasm, then wire up the front-end to run live queries against that copy—so new slices and pivots stay fast and local.
Modern business intelligence tools like Evidence use DuckDB-Wasm to bring this kind of instant interactivity to the browser. There are practical limitations, of course, but who wants a Prius for a dashboard when you could drive a Ferrari?
Running as a cloud data warehouse#
If DuckDB is so fast and powerful, why not use it as a data warehouse? You know data warehouses; they store and process all of your data in the cloud to do the single most important thing in the modern workplace: generating insights.
Turns out, several companies thought this was a clever idea. My former employer, MotherDuck, is one of them (disclaimer). By building on DuckDB as the core engine of the cloud data warehouse, MotherDuck can do some very novel things with an incredibly old category of software.
One of these is the hypertenancy architecture. Instead of many users sharing compute resources (“multitenancy”), MotherDuck users (or their agents) get their own dedicated DuckDB compute nodes. In short, this means that many users can query the same data without contending for compute resources. This model makes for superbly fast, cost-efficient analytics.
Speaking of animal-themed software companies, PostHog recently shared their migration from ClickHouse (another analytical database) to DuckDB for their data warehouse product.
Challenges with ClickHouse’s tooling ecosystem and sharing resources across many users led them to implement DuckDB, offering their end users a much friendlier, more extensible data warehouse. They even built DuckHog, a DuckDB extension that allows users to query data in the cloud data warehouse using compute power from their local machine.
PostHog implemented DuckDB without “forking the codebase or painful UDF deployments.” Translated: this thing is way easier to build and maintain.
When such vaunted, zoological companies bet the farm on DuckDB, you can be sure there’s something worth quacking about (alright, that’s it). And if you don’t believe me, just ask your neighborhood coding agent.