Breakpoint

Apple Machine Learning Research engineering blog

Limits of Confidence in Diffusion
The paper analyzes when discrete diffusion samplers can reproduce their training distribution while writing multiple token positions per step. It shows that per-position confidence scores cannot capture dependencies among jointly written tokens, and validates the resulting distributional error on the ScanAndAdd task.
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
The paper evaluates whether elaborate multi-agent harnesses improve autonomous machine learning engineering agents under equal time budgets and identical frontier models. Large-scale ablations find that a minimal coding-agent setup with read, write, and bash access performs comparably to open-source state-of-the-art harnesses, suggesting model capability is the primary driver.
How to Guide Your Language Flow
Apple researchers introduce probe guidance, a method for guiding flow-matching and diffusion language models using frozen internal states from an existing model. The approach avoids an additional inference-time forward pass, improves unconditional generation and question-answering results, and provides insight into how autoguidance works.
Dynamically Scaled Activation Steering
Apple researchers introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework that adaptively adjusts steering strength across inputs and model layers. The approach improves the trade-off between toxicity mitigation and utility, extends to text-to-image diffusion models, and adds minimal computational overhead.
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
Apple researchers introduce REVERSAL-BENCH, a benchmark that varies environmental reversibility and provides a reset oracle for evaluating reset-free reinforcement learning. Experiments across eight manipulation settings and five physics engines reveal a sharp failure cliff: agents become trapped in irrecoverable states as environments grow less reversible, while episodic agents continue learning.
How Value Induction Reshapes LLM Behaviour
Apple researchers investigate how fine-tuning language models on curated value-oriented preference data changes behaviour beyond the targeted trait. Their experiments show that value induction can affect related and contrasting values, improve safety, and increase anthropomorphic, validating, and sycophantic language.
Shared Selective Persistent Memory for Agentic LLM Systems
Apple researchers introduce shared selective persistent memory for agentic LLM systems, retaining reusable task specifications, schemas, tool configurations, and output constraints while discarding stale reasoning traces. Their deployed architecture improves task completion to 96%, cuts token costs by 97×, and enables zero-token data refreshes with git-backed artifact versioning.
Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs
Glyph is a production agentic system for generating column descriptions and assigning sensitivity labels in enterprise data catalogs. It combines code-grounded retrieval, parallel tagging strategies, a fine-tuned MiniLM encoder, and Reciprocal Rank Fusion to deliver auditable, value-free classification with measurable retrieval improvements.
SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign
Apple researchers introduce SimpleDesign, an end-to-end multimodal generative model for jointly designing protein sequences and three-dimensional structures. The Transformer-based approach combines sequence cross-entropy with structural regression, avoids separate autoencoder pretraining, and achieves competitive results on co-design and unconditional generation benchmarks using more than 2 million sequence–structure pairs.