Apple researchers investigate why multilingual self-supervised speech models lag behind monolingual models under matched data budgets. Using English/French HuBERT experiments, they show that auxiliary language classification and per-language clustering improve phonetic, lexical, and prosodic performance while retaining cross-language sharing.
Apple Machine Learning Research engineering blog
The paper analyzes when discrete diffusion samplers can reproduce their training distribution while writing multiple token positions per step. It shows that per-position confidence scores cannot capture dependencies among jointly written tokens, and validates the resulting distributional error on the ScanAndAdd task.
The paper evaluates whether elaborate multi-agent harnesses improve autonomous machine learning engineering agents under equal time budgets and identical frontier models. Large-scale ablations find that a minimal coding-agent setup with read, write, and bash access performs comparably to open-source state-of-the-art harnesses, suggesting model capability is the primary driver.
Apple researchers introduce RLTL;DR, a reinforcement-learning method that has models generate and internalize concise feedback after failed attempts. Experiments show substantial gains on difficult coding and tool-calling tasks, with SFTL;DR demonstrating that compact task-to-insight training can recover most of the performance.
Apple researchers systematically evaluate LLM conditioning methods for injecting and removing concepts, measuring both effectiveness and fluency. The study finds that activation steering can significantly harm fluency and is less effective on instruction-tuned models, while prompting and supervised fine-tuning perform differently across injection and removal tasks.
Apple researchers introduce SCLATE, an execution substrate that coordinates benchmark and agent-side events on a shared hybrid clock for continual-learning evaluation and training. The system supports unmodified agent harnesses and memory, records model-call data, and demonstrates measurable gains in file efficiency, SWE-bench performance, and held-out accuracy.
Apple researchers introduce a round-trip protocol for measuring how much tree-structured information language models preserve when serialized as natural language. Evaluating 16 models reveals asymmetric communication quality, generation as the main failure source, and substantial gains from targeted fine-tuning.
The paper develops improved convergence guarantees for federated stochastic variational inequalities, including a refined analysis of Local Extra SGD. It introduces LIPPAX to reduce client drift and extends the results to bounded-Hessian, bounded-operator, low-variance, and composite settings.
Apple researchers present a practical recipe for semi-supervised federated ASR using online pseudo-labeling and server-side stabilization. They show how teacher selection, labeled-data anchoring, augmentation, batch size, and domain overlap affect convergence, reporting gains over prior methods across in-domain and cross-domain evaluations.
Apple researchers present a latent-space distillation method for compressing streaming neural audio tokenizers used in on-device Dictation. By training a smaller encoder to match the teacher’s pre-quantizer representations, they achieve 2.8× compression while keeping word error rate within 1.9% on most evaluated teacher–student pairs.
Apple researchers introduce probe guidance, a method for guiding flow-matching and diffusion language models using frozen internal states from an existing model. The approach avoids an additional inference-time forward pass, improves unconditional generation and question-answering results, and provides insight into how autoguidance works.
Apple researchers introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework that adaptively adjusts steering strength across inputs and model layers. The approach improves the trade-off between toxicity mitigation and utility, extends to text-to-image diffusion models, and adds minimal computational overhead.
Apple researchers introduce REVERSAL-BENCH, a benchmark that varies environmental reversibility and provides a reset oracle for evaluating reset-free reinforcement learning. Experiments across eight manipulation settings and five physics engines reveal a sharp failure cliff: agents become trapped in irrecoverable states as environments grow less reversible, while episodic agents continue learning.
Apple researchers introduce Trajectory-Shaped Discrete Flow Matching, which uses an energy-based compass during training to replace blind stochastic jumps with more coherent generation paths. An 8-step, 170M-parameter student achieves 32% lower perplexity than a 1,024-step teacher while running 128× faster.
Apple researchers investigate how fine-tuning language models on curated value-oriented preference data changes behaviour beyond the targeted trait. Their experiments show that value induction can affect related and contrasting values, improve safety, and increase anthropomorphic, validating, and sycophantic language.
Apple researchers introduce DACA-GRPO, a plug-in enhancement for reinforcement learning in diffusion language models. It assigns credit across denoising steps and reduces likelihood-estimation bias through progress-based token weighting and stratified masking, improving results across reasoning, code, constraint, and structured-generation benchmarks.
Apple researchers introduce shared selective persistent memory for agentic LLM systems, retaining reusable task specifications, schemas, tool configurations, and output constraints while discarding stale reasoning traces. Their deployed architecture improves task completion to 96%, cuts token costs by 97×, and enables zero-token data refreshes with git-backed artifact versioning.
Glyph is a production agentic system for generating column descriptions and assigning sensitivity labels in enterprise data catalogs. It combines code-grounded retrieval, parallel tagging strategies, a fine-tuned MiniLM encoder, and Reciprocal Rank Fusion to deliver auditable, value-free classification with measurable retrieval improvements.
Apple researchers introduce SimpleDesign, an end-to-end multimodal generative model for jointly designing protein sequences and three-dimensional structures. The Transformer-based approach combines sequence cross-entropy with structural regression, avoids separate autoencoder pretraining, and achieves competitive results on co-design and unconditional generation benchmarks using more than 2 million sequence–structure pairs.
DiscoSign presents an LLM-based framework for translating English text into discourse-aware American Sign Language gloss. It addresses spatial coreference, question-answer clauses, and concept-gloss consistency, introducing dedicated evaluation metrics and demonstrating improved entity tracking and spatial consistency over sentence-level translation.