Breakpoint

DuckDB engineering blog

Faster String Aggregations with Dimension Tables
The post explains how encoding repeated, long strings as sorted, narrow integer keys in dimension tables can make DuckDB aggregations faster and more memory-efficient. It covers schema construction, query patterns, ENUM alternatives, maintenance trade-offs, benchmarking, and cases where encoding does not help.
Jev and DuckDB: Plain-English Conditions in SQL
The post explains how Jev brings typed, plain-English judgments to DuckDB, enabling SQL filtering, classification, ranking, and scoring over semi-structured data. It compares community extensions, batching and caching strategies, performance trade-offs, costs, and privacy considerations.
Announcing DuckDB 1.5.6
DuckDB 1.5.6 is a patch release containing correctness fixes, crash and recovery fixes, security hardening, and C API improvements. The post also previews DuckDB 2.0 performance on Windows, reporting a sixfold speedup on a TPC-H benchmark through compiler and allocator changes.
DuckDB and Hugging Face: Querying Datasets Directly
The post explains how DuckDB’s hf:// protocol queries Hugging Face datasets remotely without first downloading them. It demonstrates globbing files, pinning revisions, accessing private datasets, filtering into local Parquet files, and joining Hub data with local tables.
DuckDB Now Ships inside dbt v2
The post explains how dbt v2 integrates DuckDB through its Rust-based Fusion engine, eliminating the separate adapter installation. It covers setup, DuckLake and Iceberg catalogs, querying dbt metadata as Parquet, column-level lineage, pinned DuckDB versions, and migration from dbt v1.
Persistent Databases in the Browser with DuckDB-Wasm and OPFS
The post explains how to use DuckDB-Wasm with the browser’s Origin Private File System (OPFS) for persistent databases and data files. It covers version pitfalls, WAL and checkpoint durability, automatic versus manual file handling, caching remote data, and exporting databases or Parquet files.
DuckDB Skills for Claude Code
DuckDB’s duckdb-skills plugin lets Claude Code query, profile, convert, and explore local and remote data through the DuckDB CLI instead of ad hoc Python scripts. The post explains its skills, SQL-based session state, error-retry workflow, and support for cloud storage and spatial data.
Try DuckDB v2.0-dev
DuckDB announces the v2.0-cyanoptera feature freeze and invites users to test alpha clients, extensions, and existing SQL workloads ahead of the planned October release. The post provides installation commands for the CLI and Python client, outlines extension changes, and explains how to report reproducible issues.
DuckLabs to Join AWS, Projects to Remain Open Source
DuckLabs announces that it will join Amazon Web Services while DuckDB, DuckLake, Quack, and related projects remain open source under the MIT license. The projects will continue under the DuckDB Foundation with unchanged roadmaps and governance.
How DuckDB Runs Recursive CTEs Faster
DuckDB’s upcoming v2.0 accelerates recursive CTEs by retaining epoch-invariant state, adapting execution to exact frontier cardinalities, and probing keyed state directly. The post explains the execution architecture, repeatability safeguards, changed-key semantics for USING KEY ... UNION, and measured speedups across graph and stateful workloads.