Fly.io introduces MCP support for Sprites, disposable cloud computers with durable filesystems that agents can provision and control. The post explains progressive capability disclosure, MCP resources and safety annotations, scoped authentication, and practical workflows for testing, benchmarking, and running services.
Fly.io engineering blog
Fly.io outlines its strategic shift toward Sprites, cloud computers designed for coding agents, and explains their durable storage, rapid cloning, metered execution, and authenticated Connectors. The post also announces new funding and a transition to Scott Johnston as CEO.
The post explains how to isolate AI-agent execution from the agent process using disposable, per-session or per-task sandboxes. It covers ephemeral credential injection, checkpoint-based rollback, lifecycle trade-offs, and why separating an agent’s durable home from its untrusted command environment improves safety and reliability.
Ben Johnson explains how Litestream’s writable VFS enables SQLite to serve reads and buffered writes directly from S3-compatible object storage during fast Sprite cold starts. The post covers single-writer constraints, eventual durability, page indexing, LTX synchronization, and background database hydration.
Fly.io explains how Sprites provide instant, persistent, auto-sleeping Linux virtual machines. The post details their preallocated runtime model, object-storage-backed filesystems, checkpointing, and inside-out orchestration, along with the trade-offs versus container-based Fly Machines.
Fly.io argues that coding agents need durable, disposable cloud computers rather than short-lived, read-only sandboxes. It introduces Sprites, which provide fast-booting Linux environments with persistent storage, automatic hibernation, HTTPS networking, and rapid checkpoint and restore.
Litestream VFS lets SQLite query backups directly from object storage without downloading an entire database. The post explains how LTX compaction, indexed range reads, point-in-time recovery, and an LRU cache enable fast, near-real-time remote replicas and efficient historical queries.
Thomas Ptacek demonstrates how a few lines of Python and the OpenAI Responses API can become a tool-using LLM agent. He explains stateless model calls, context engineering, sub-agents, tool design, security boundaries, and the open engineering trade-offs involved in building reliable agent systems.
Fly.io explains Corrosion, its open-source Rust service for synchronizing SQLite state across globally distributed workers without centralized consensus. The post examines outages caused by deadlocks, schema backfills, and cascading updates, then details watchdogs, CRDTs, checkpointing, and regionalization used to reduce failure blast radius.
Fly.io details how its CEO was phished through a fake X.com login, leading to an account takeover and crypto scam posts. The incident explains why phishing-resistant MFA, passkeys, SSO, credential hygiene, and rapid access auditing are more reliable than relying on users not to click malicious links.
Litestream v0.5.0 replaces WAL segments with the transaction-aware LTX format, enabling hierarchical compaction and efficient point-in-time recovery for SQLite. The release also removes generations, adds NATS JetStream support, improves storage backends, and simplifies cross-compilation by eliminating CGO.
MorphLLM’s Fast Apply API enables AI coding agents to make precise, structure-aware edits without rewriting entire files. The post explains its integration workflow and reports benchmark results comparing its speed and accuracy with traditional search-and-replace approaches.
The post explains how AI software builders can calibrate user trust to match system capabilities, reducing both over-reliance and underuse. It presents practical design strategies involving cooperative versus delegative systems, adaptive feedback, layered explanations, capability boundaries, and trust-calibration metrics.
The post argues that games and simulated environments provide more realistic signals for evaluating AI models than static benchmarks, testing strategic reasoning, memory, adaptation, and conversational behavior. It also introduces a one-click Fly.io deployment of AI Town for comparing OpenAI-compatible services, Together.ai models, and custom embeddings with scale-to-zero cost optimization.
The post argues that AI products should prioritize deep specialization in a chosen model over broad model agnosticism, because prompts, behavior, workflows, and user trust are tightly coupled. It also emphasizes treating model evaluation as an architectural concern and suggests game-like environments for testing model behavior.
Chris McCord introduces Phoenix.new, a browser-based AI coding agent running in isolated Fly Machines. He explains how root-level VM access, headless browser testing, live previews, cloud deployment, and GitHub integration let agents build, test, and deploy real-time Phoenix applications.
The post explains the Model Context Protocol through comparisons with Alexa Skills, APIs, and API introspection. It also examines MCP server lifecycles, tool design, security risks, remote deployment, and the potential for MCPs to enable personalized AI agents.
The author argues that modern coding agents already automate much of software development’s tedious work by navigating codebases, editing files, running tools, and iterating against tests. The essay explains why engineers should focus on judgment, curation, and guardrails while treating LLMs as powerful assistants rather than autonomous replacements.
The post introduces a practical, opinionated guide to deploying real applications with Kamal 2.0 and Docker. It covers supporting production concerns including container builds and registries, secrets, hosting, load balancing, databases, backups, searchable logs, monitoring, and security.
Fly.io investigates a severe Anycast routing outage caused by a subtle parking_lot RWLock corruption bug in its Rust proxy. The post traces the debugging process from deadlock symptoms and core dumps to the faulty timeout wake-up path, and details the instrumentation and locking changes that improved resilience.