Breakpoint

Shopify engineering blog

ShopGym: Realistic, reproducible sandboxes for shopping agents
ShopGym converts live storefronts into anonymized, resettable sandbox shops and generates grounded shopping tasks for agent evaluation. The post explains its exploration, staged generation, verification, and benchmarking workflow, including structural and behavioral comparisons across real and synthetic stores.
Helix: The internal tool powering our Shopify app's native migration
Shopify explains how Helix uses LLM agents, incremental checkpoints, behavioral tests, visual comparison, adversarial code reviews, and engineer approval to migrate a 300-screen React Native app to native Swift and Kotlin. The checkpoint-and-gate workflow prioritizes reliable convergence over perfect first attempts while preserving code quality and UI fidelity.
Migrating Shop app from React Native to native
Shopify details how it rebuilt the Shop app from React Native in Swift and Kotlin in 12 weeks with coding agents. The migration improved startup time, Android app size, build times, rendering performance, and session stability while preserving feature parity, analytics, authentication, and notifications.
Native is now the future of mobile at Shopify
Shopify explains why improved coding agents changed the trade-offs behind its mobile stack, prompting a move from React Native back to Swift and Kotlin. The post details its greenfield migration strategy, Helix’s checkpoint-based review system, and CLI-driven architecture for faster agent feedback loops.
How River takes security work from a fix to merge
Shopify explains how River, its Slack-based AI agent, moves vulnerability remediation from detection and patch creation through rebasing, CI validation, human handoff, merge, and verified closure. The workflow revalidates live repository state, binds evidence to current commits, preserves investigation context, and uses deterministic controls for security-critical guarantees.
Teaching Sidekick to say no: automated data curation with LLM judge consensus
Shopify describes how it improved Sidekick’s customer segmentation skill by teaching the model when to refuse impossible requests. The post details an automated data curation pipeline that uses a calibrated ensemble of LLM judges, strict consensus, and a mutually exclusive refusal taxonomy to resolve conflicting labels and create higher-quality training data. The result was better refusal behavior, more stable training, and measurable gains over naive dataset merging.
Under the River
What it took to ship Shopify's Slack-native agent River, including lessons learned and the underlying substrate that powers it. The post is co-authored by River.
Shopify’s journey to faster breadth-first GraphQL execution
Shopify analyzed hidden CPU and memory costs in conventional depth-first GraphQL execution for high-cardinality, deeply nested queries and implemented GraphQL Cardinal, a breadth-first execution engine that resolves each field once across aggregated object sets. Cardinal delivered up to ~15x faster CPU-bound field execution and ~90% less memory in large-list tests; the post also details migration strategies (an interpreter for legacy resolvers, tracer adaptations, and incremental breadth-style resolver rewrites) and tradeoffs such as error-reporting behavior and enqueuing-driven execution.