Google Research presents a TEE-based federated learning system that provides externally verifiable privacy guarantees through encrypted uploads, access policies, remote attestation, differential privacy, and reproducible builds. The architecture shifts training computation to servers, improving device coverage, training speed, and model accuracy while supporting fault-tolerant recovery.
Google Research engineering blog
Google Research introduces Diffusion Controller, a lightweight control network that steers image-generation models toward better prompt and preference alignment while keeping the base model frozen. The framework unifies inference-time guidance and fine-tuning, supports restricted-access models, and uses PPO or reward-weighted loss to improve quality without destabilizing generation.
Google researchers present a unified multi-agent framework for generating coherent, minutes-long video narratives. Its components use hierarchical planning, persistent visual memory, adaptive interpolation and extrapolation, and closed-loop prompt refinement to reduce character drift, scene inconsistency, and cascading pipeline failures.
Google Research introduces MilleMiglia, an open-source C++ generator for realistic, privacy-preserving middle-mile logistics benchmarks. It models space-time flows, fixed vehicle schedules, distribution-center capacity, shipment synchronization, and scalable instances for optimization and machine-learning research.
Google Research presents a generative UI system that helps teachers create curriculum-aligned, interactive STEM simulations. The approach combines structured learning objectives, progressively difficult game levels, AI-generated scaffolding, iterative agentic evaluation, and teacher validation to improve educational quality and solvability.
Google Research presents Retrieve-for-Train, a framework that uses offline reinforcement learning to compile set-level search behavior into training data for a lightweight diffusion retriever. The approach generates diverse, grounded query fan-outs in a single non-autoregressive pass, achieving reported 12–20× inference speedups while avoiding costly test-time reasoning.
ToolGrad introduces an answer-first framework for generating tool-use datasets by constructing verified API workflows before synthesizing user prompts. Its proposer, executor, selector, and textual-gradient updater modules produce higher-pass-rate, lower-cost training data and improve Gemma models on out-of-distribution function-calling benchmarks.
Google Research details a complete, human-verified connectome of the male fruit fly, containing more than 166,000 neurons and 125 million synaptic connections. The post explains how AI-based image reconstruction, flood-filling networks, and the PATHFINDER system support large-scale brain mapping and open research datasets.
The study evaluates transfer learning for polygenic risk scores across European and Japanese cohorts, showing that European data improves prediction when target populations are small but can reduce accuracy as local samples grow. It compares elastic net, meta-analysis, and PRS-CSx approaches and identifies trait-dependent sample-size thresholds for effective cross-population modeling.
Google Research presents MAPL-EMIT, a Swin-S vision transformer that detects, quantifies, and localizes methane plumes in hyperspectral satellite data. The post explains its physics-based synthetic plume training approach, performance against expert annotations, handling of overlapping emissions, and released datasets and inference tools.
Google Research introduces TimesFM-3, a 330-million-parameter time-series foundation model for zero-shot multivariate forecasting. Its alternating temporal and cross-series attention, covariate lookahead, and non-autoregressive decoding generate probabilistic forecasts in a single pass, achieving leading results across three public benchmarks.
Google Research presents the Planetary Prediction Engine, an autonomous system that discovers geospatial data, curates multimodal features, trains models, and evaluates predictions from natural-language queries. Its leakage controls, overfitting safeguards, and benchmark results demonstrate improved performance across public health, food security, environmental risk, and outbreak nowcasting.
Google Research introduces GlucoFM, a lightweight self-supervised foundation model for continuous glucose monitoring. Its dual-stream architecture separates slow glycemic trends from short-term deviations, improving metabolic risk assessment, postprandial response forecasting, cross-cohort transfer, and few-shot learning.
Google researchers present AgentHands, an LLM-powered XR system that converts conversational responses into spatially anchored, co-speech hand gestures. The prototype combines scene registration, gesture-event generation, word-level TTS synchronization, and user studies showing improved spatial grounding, action comprehension, and safety-cue recognition.
Google Research introduces Mobility-Embedded POIs (ME-POIs), a framework that combines language-based place descriptions with aggregated, anonymized mobility patterns. Its multiscale propagation and text-mobility alignment improve predictions of visit intent, price level, opening hours, closures, and busyness on unseen places.
Google Research presents the Biomarker Discovery Framework, a human-supervised multi-agent system that combines deterministic statistical analysis, machine learning, literature grounding, and adversarial validation to prioritize wearable-derived biomarkers. Evaluations across 9,279 participant-observations show convergent candidate signals and improved downstream prediction while emphasizing leakage controls, uncertainty, and non-causal interpretation.
Google Research presents PhotoScan, a deep learning framework that estimates body-fat percentage and regional fat ratios from standard smartphone photos. Trained on UK Biobank data and validated against DXA scans, the system improved insulin-resistance prediction to near-DXA performance in clinical research cohorts.
This post presents a knowledge profiling framework for diagnosing factual errors in large language models by separating encoding from recall. Using the WikiProfile benchmark and evaluations across 13 models, the authors find that frontier LLMs encode most facts but still struggle to retrieve many of them directly, especially for rare facts and reverse questions. The post also shows that thinking improves recovery of encoded-but-inaccessible facts, suggesting factuality gains may come more from better utilization than from scaling alone.
This post introduces AMIE (Video), an audio-visual clinical AI system built on Gemini and Project Astra that can conduct real-time synchronous video consultations. The research reports expert-level performance in a randomized OSCE-style study, with strong results in clinical reasoning, physical examination guidance, and patient-actor preference compared with text-only and physician baselines.
Google Research introduces the Science One Framework, an experimental autonomous research prototype designed to make AI-generated scientific work verifiable by construction. The post explains Chain-of-Evidence, a framework that binds claims to evidence, and CoE Audit, an automated evaluation protocol that checks score integrity, reference validity, and alignment between methods and code.