Breakpoint

PlanetScale Engineering engineering blog

Designing Neki for performance
The post explains how Neki improves sharded PostgreSQL performance by preserving wire-format messages and decoding data lazily. It details composable representations for routing, limits, sorting, projection, joins, and aggregation, showing how byte-level reuse avoids unnecessary allocations, parsing, and serialization.
Handling hot shards
The post explains how tenant growth and changing access patterns can make an initially sensible shard key produce hot shards. It compares vertical scaling and tenant isolation with table-specific resharding, showing how declarative topology and online data movement can redistribute load without downtime.
When to choose x86-64 vs aarch64
Ahmed Darwich compares ARM and x86-64 cloud CPUs for PostgreSQL workloads, explaining vCPU differences, parallelism, clock speed, vector instructions, and extension compatibility. He outlines when each architecture is preferable and how to benchmark or migrate between them safely.
TIN Postgres search is faster, better, and cheaper
Patrick Reynolds demonstrates how TIN accelerates Postgres full-text search with immediate keystroke queries, wildcard autocomplete, fuzzy matching, stop-word support, and exact COUNT(*) results. The comparison reports 11ms ranking, 213,447 exact matches versus an estimated 3,400, and substantially lower infrastructure costs.
Anatomy of a (Postgres) search engine
The post explains how full-text search engines use inverted indexes, postings lists, positional and frequency data, BM25 scoring, and segmented storage. It then details the additional challenges of implementing these structures as a transactional, mutable Postgres index, including visibility, VACUUM, tombstones, merging, and query planning.
Blocking cutovers to save replication slots
Simeon Griggs explains how PostgreSQL logical replication slots can become invalid during a replica promotion, causing downstream CDC consumers to miss changes. The post details the failover, hot_standby_feedback, and sync_replication_slots settings, along with PlanetScale’s safeguards for blocking unsafe cutovers.
The architecture of Neki
PlanetScale explains Neki’s architecture for sharding ordinary PostgreSQL across a single client-facing endpoint. The post details its Router, Sidecar, Admin, Operator, topology, and Replicator components, including failover, distributed query execution, OID consistency, resharding, and online schema changes.
Introducing Lead: TIN-compatible full-text search for CI
PlanetScale introduces Lead, a deliberately slow, TIN-compatible PostgreSQL full-text search extension for CI, development, and staging. It preserves TINQL queries, tokenization, BM25 scoring, highlighting, transactional visibility, and durability while using full table scans, making it practical for small test datasets rather than production workloads.
Introducing TIN: full-text search for Postgres
PlanetScale introduces TIN, a full-text search index for Postgres supporting Boolean, phrase, fuzzy, wildcard, regex, COUNT, and BM25 queries. The post details its benchmark results and explains how ctid-based bitmap indexing, vectorization, MVCC handling, and segment merging deliver substantially higher performance than competing indexes.
118 million queries per second on Neki
PlanetScale reports sustaining 118.5 million read queries per second on Neki, its sharded Postgres platform, across 512 shards holding 1.22 PiB of data. The benchmark demonstrates near-linear scaling, with detailed throughput, latency, error-rate, IOPS, and network results.
The lifecycle of a sharded Postgres query
PlanetScale traces a sharded Postgres query from authentication and wire-protocol handling through distributed planning, cross-shard joins, connection pooling, and result assembly. The post explains routing strategies, hash joins, memory spilling, sidecars, and router-level execution trade-offs.
What is a Neki router?
The post explains how Neki routes application queries across sharded PostgreSQL databases while presenting a standard Postgres connection. It covers stateless router fleets, sidecars, connection scaling, and the distinction between Neki’s shard-selection plan and PostgreSQL’s per-shard execution plan.
Problems with large tables in Postgres
Simeon Griggs explains how large, wide, and bloated PostgreSQL tables create vacuum, query, backup, indexing, and replication problems. The post compares partitioning, vertical scaling, and sharding, detailing their trade-offs and how sharding isolates workloads and cluster limits.
The history of Postgres sharding
The post traces two decades of Postgres sharding, from Skype’s PL/Proxy and Instagram’s application-level approach to Citus, distributed Postgres-compatible databases, and modern router-based systems. It explains the trade-offs among explicit and automatic sharding, including query routing, cross-shard latency, operational complexity, and data placement.
Poisoned Postgres connection pools
PlanetScale explains how session-level PostgreSQL settings can poison PgBouncer transaction-mode connection pools and make applications appear read-only. It details how to diagnose the failure, recover connections with DISCARD ALL, and prevent leaks through transaction-scoped settings, replica routing, and safer cleanup.
What is a data topology?
PlanetScale explains how Neki uses data topologies to map logical PostgreSQL tables to physical shards and route queries. The article details shard indexes, shard groups, routing strategies, colocation, inheritance, and how topology changes support resharding without interrupting traffic.
The dangers of Postgres subtransactions
PostgreSQL subtransaction cache overflow can cause cluster-wide performance cliffs by forcing expensive pg_subtrans lookups and SLRU lock contention. The post explains how overflowed transaction state delays hot standby on new replicas and presents benchmarks, detection queries, and operational mitigations.
Massively parallel Postgres backups
PlanetScale explains how its Neki sharded Postgres platform creates consistent, encrypted backups for petabyte-scale databases. The design uses temporary per-shard compute, object storage, WAL replay, and hybrid replication to achieve high throughput while minimizing production impact.
Postgres backups under the hood
The post explains PostgreSQL’s logical, filesystem, and continuous-archiving backup methods, including how MVCC snapshots can trigger transaction wraparound on large databases. It details WAL, full-page writes, point-in-time recovery, and PlanetScale’s shard-based approach to keeping backups and restores practical at scale.