Grab presents a feedback-driven framework for validating knowledge-graph relationships through live search interactions. It combines candidate-edge injection, exploration–exploitation, weighted behavioral signals, confidence tiers, and anti-abuse controls to promote accurate links and prune AI-generated errors.
Grab Tech engineering blog
Grab explains how Jarvis Pro routes account-manager requests before generating answers, using task-specific context, narrow memory, authorization checks, and guardrails. The post details metric reconciliation, latency trade-offs, and layered offline evaluations that improved routing and answer quality while avoiding unsafe recommendations.
Grab describes how it used agentic AI, medium tests, and a patch-test-compare workflow to migrate hundreds of services to Distroless container images safely. The approach combines MCP integrations, repository-specific skills, automated CI remediation, batch changes, and human review to reduce migration toil while catching runtime dependency regressions.
Grab describes how agentic analytics workflows are moving analysts from manually producing reports to owning questions, judgment, and decisions. The post details its autonomy ladder, certified context, routing frameworks, automated root-cause analysis, self-healing pipelines, and measured improvements in cycle time and self-service adoption.
Grab details the zero-downtime migration of its high-volume fraud-detection Counter Service from a wide-column database to Aerospike. The post explains the Rust storage abstraction, shadow-read rollout, map-based data model, atomic updates, client pitfalls, and resulting improvements in latency, storage, and cost.
Grab details Palana, a Kubernetes-native platform for securely running autonomous AI agents. The post explains its per-agent isolation, identity and least-privilege secret model, mediated network and LLM access, observability, and lifecycle controls for attributable and recoverable agent operations.
Grab explains how Grab Bench evaluates AI systems on production-shaped tasks without exposing live data. The configurable harness uses task-specific contracts, deterministic or LLM-based scoring, hidden tests, baselines, canaries, and row-level failure records to expose plausible but incorrect behavior.
Grab explains how an internal support bot evolved into LLM-Kit, a framework for building and operating AI agents in production. The post details its evaluation scaffolding, model gateway integration, observability, MCP tool discovery, secrets management, and gRPC service patterns, showing how reusable infrastructure reduced agent setup time from weeks to about an hour.
Grab details its migration from Hive Parquet to Apache Iceberg across a petabyte-scale data lake, addressing metadata latency, small files, operational overhead, and data consistency. It also explains the design of UnifiedSparkCatalog and reports substantial gains in query performance, S3 API costs, and compute efficiency.
Grab explains how automated Data Production Issues operationalize data reliability across its data mesh. The system triages contract breaches, performs cross-platform root-cause analysis, routes ownership, and safely auto-heals routine failures, reducing manual effort and accelerating recovery.