ShopGym converts live storefronts into anonymized, resettable sandbox shops and generates grounded shopping tasks for agent evaluation. The post explains its exploration, staged generation, verification, and benchmarking workflow, including structural and behavioral comparisons across real and synthetic stores.
Shopify engineering blog
Shopify explains how Helix uses LLM agents, incremental checkpoints, behavioral tests, visual comparison, adversarial code reviews, and engineer approval to migrate a 300-screen React Native app to native Swift and Kotlin. The checkpoint-and-gate workflow prioritizes reliable convergence over perfect first attempts while preserving code quality and UI fidelity.
Shopify details how it rebuilt the Shop app from React Native in Swift and Kotlin in 12 weeks with coding agents. The migration improved startup time, Android app size, build times, rendering performance, and session stability while preserving feature parity, analytics, authentication, and notifications.
Shopify explains why improved coding agents changed the trade-offs behind its mobile stack, prompting a move from React Native back to Swift and Kotlin. The post details its greenfield migration strategy, Helix’s checkpoint-based review system, and CLI-driven architecture for faster agent feedback loops.
Shopify explains how River, its Slack-based AI agent, moves vulnerability remediation from detection and patch creation through rebasing, CI validation, human handoff, merge, and verified closure. The workflow revalidates live repository state, binds evidence to current commits, preserves investigation context, and uses deterministic controls for security-critical guarantees.
Based on the title alone, the post presents gisting as a technique for compressing LLM-agent context to increase throughput and reduce inference costs.
We rebuilt our mobile end-to-end testing framework with a strict API and computer vision, raising test stability drastically.
How we compress production failures into model weights every day, beat frontier-model quality, and cut serving costs 96%.
We built an agentic code review and test oracle harness that discovers vulnerabilities, proves them with real tests, and provides Shopify-tuned fixes.
We moved five high-traffic checkout extensions to remote-dom and Polaris web components, cutting bundle sizes drastically and making checkout faster.
How a Hack Days team built a music player, custom GLSL visualizers, and an artist toolkit for storefronts, all in three days.
Building an LLM-powered pipeline to match product listings across millions of merchants into a unified catalog that AI agents can search.
Shopify describes how it improved Sidekick’s customer segmentation skill by teaching the model when to refuse impossible requests. The post details an automated data curation pipeline that uses a calibrated ensemble of LLM judges, strict consensus, and a mutually exclusive refusal taxonomy to resolve conflicting labels and create higher-quality training data. The result was better refusal behavior, more stable training, and measurable gains over naive dataset merging.
Quick lets anyone at Shopify ship a site in seconds. It has changed the culture of how we build and share.
What it took to ship Shopify's Slack-native agent River, including lessons learned and the underlying substrate that powers it. The post is co-authored by River.
How we used SKIP LOCKED, composite primary keys, and connection visibility to hit our scale targets.
We fine-tuned Qwen3-32B into a tool-calling agent that generates Flow automations from natural language. The post highlights how the system became faster, cheaper, and more accurate than the frontier model it replaced, with a weekly retraining flywheel based on real merchant data.
Tobi and I generalized Karpathy’s Autoresearch to improve 40+ metrics across Shopify, then we open-sourced our project.
The technical blueprint for an AI-powered mirror that sees customers, analyzes their appearance, and delivers personalized product recommendations in real-time.
Shopify analyzed hidden CPU and memory costs in conventional depth-first GraphQL execution for high-cardinality, deeply nested queries and implemented GraphQL Cardinal, a breadth-first execution engine that resolves each field once across aggregated object sets. Cardinal delivered up to ~15x faster CPU-bound field execution and ~90% less memory in large-list tests; the post also details migration strategies (an interpreter for legacy resolvers, tracer adaptations, and incremental breadth-style resolver rewrites) and tradeoffs such as error-reporting behavior and enqueuing-driven execution.