Breakpoint

Grab Tech engineering blog

Building Jarvis Pro: Route first, answer later
Grab explains how Jarvis Pro routes account-manager requests before generating answers, using task-specific context, narrow memory, authorization checks, and guardrails. The post details metric reconciliation, latency trade-offs, and layered offline evaluations that improved routing and answer quality while avoiding unsafe recommendations.
Scaling out Distroless adoption With AI
Grab describes how it used agentic AI, medium tests, and a patch-test-compare workflow to migrate hundreds of services to Distroless container images safely. The approach combines MCP integrations, repository-specific skills, automated CI remediation, batch changes, and human review to reduce migration toil while catching runtime dependency regressions.
How AI is transforming analytics at Grab
Grab describes how agentic analytics workflows are moving analysts from manually producing reports to owning questions, judgment, and decisions. The post details its autonomy ladder, certified context, routing frameworks, automated root-cause analysis, self-healing pipelines, and measured improvements in cycle time and self-service adoption.
Migrating Counter Service storage: Design choices and learnings
Grab details the zero-downtime migration of its high-volume fraud-detection Counter Service from a wide-column database to Aerospike. The post explains the Rust storage abstraction, shadow-read rollout, map-based data model, atomic updates, client pitfalls, and resulting improvements in latency, storage, and cost.
Grab Bench: Evaluating AI on Grab-shaped production work
Grab explains how Grab Bench evaluates AI systems on production-shaped tasks without exposing live data. The configurable harness uses task-specific contracts, deterministic or LLM-based scoring, hidden tests, baselines, canaries, and row-level failure records to expose plausible but incorrect behavior.
Agent platform (Part 1): How we help Grab build and run AI agents at scale
Grab explains how an internal support bot evolved into LLM-Kit, a framework for building and operating AI agents in production. The post details its evaluation scaffolding, model gateway integration, observability, MCP tool discovery, secrets management, and gRPC service patterns, showing how reusable infrastructure reduced agent setup time from weeks to about an hour.
Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption
Grab details its migration from Hive Parquet to Apache Iceberg across a petabyte-scale data lake, addressing metadata latency, small files, operational overhead, and data consistency. It also explains the design of UnifiedSparkCatalog and reports substantial gains in query performance, S3 API costs, and compute efficiency.