Week beginning 28th September 2026

Every computer science paper posted to arXiv this week, with plain-language summaries and practical uses for each one. Includes commercial applications where relevant.

Generating self-referential images with conformal geometric transformations

Moore, Escher, Penrose: A Conformal Golden Braid

Abstract: I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map $z \mapsto z^α$, $α\in \mathbb{C}$. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may "repair" the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse $T^\dagger$ of the non-invertible image transformation $T$, adapted to its recursive constraint. In the idealized formulation, the Penrose identity $TT^\dagger T = T$ makes $TT^\dagger$ an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with $T$ and $T^\dagger$: source-space steps develop the untwisted scene, while transformed-space steps refine its appearance and connections in the final geometry. We generate Print Gallery-like compositions and explore further transformations. Rather than distorting a finished image, we let the scene and its distortion develop together.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Some artworks show strange images that seem to look back at themselves in impossible ways. The authors studied how a famous image by M.C. Escher, called Print Gallery, uses a special kind of twisting geometry called a conformal map. They then used AI image models combined with these geometric rules to create new images that develop both the scene and its distortion together. This approach lets them produce complex, self-referential pictures that match the peculiar structure of the original artwork.
Open → 2610.02210v1

Sphere Encoder 2 improves image generation quality and speed

Sphere Encoder 2

Abstract: Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Generating images from random points inside a sphere is tricky because points tend to cluster away from the poles, creating gaps in those areas. Also, the original method tries to make images match pixel-for-pixel, leading to blurry results. The authors fixed these problems by adjusting how the model trains and where it samples points, which makes it create clearer images faster.
Open → 2610.02208v1

Gaussian blendshape distillation speeds up real-time avatar animation

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

Abstract: 3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/

Thu 1 OctComputer Vision and Pattern RecognitionArtificial IntelligenceHuman-Computer Interaction
The gist
Animating detailed 3D avatars typically requires complex neural networks that slow down real-time performance. The authors found that the complex animations of pretrained Gaussian avatars can be expressed as a simpler mix of basic facial or body shapes called blendshapes. They created a method called GALA that uses a small network to quickly predict how much to mix these blendshapes, making animation much faster without retraining existing models. This approach works well across different avatar types, allowing smooth animations even on mobile devices.
Open → 2610.02207v1

Benchmark measures language models running cybersecurity commands accurately

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Abstract: LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

Thu 1 OctComputation and LanguageArtificial IntelligenceCryptography and Security
The gist
Accurately using cybersecurity tools often means typing commands perfectly, but small mistakes can cause failure. The researchers created KaliBench, which is a big set of real-world command examples to test if language models can turn plain English into correct commands for Kali Linux tools. They tested many open models and found that all struggle to get commands exactly right. However, by training with special rewards from KaliBench, smaller models improved enough to rival much larger ones. This helps make AI better at safely and reliably assisting in cybersecurity tasks.
Open → 2610.02206v1

Programmable world model benchmark tests video model rule following

ROWBench: Do Video Models Render What the Program Specifies?

Abstract: Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
It is hard to tell if video-based AI models truly understand and follow all the details of a programmed world. The authors created PROWBench, a set of many short videos with detailed records of what should happen in the scene. This lets them check if a video model’s outputs really match the program’s rules and events, even ones outside the camera view. They also built tools to generate and view these scenes from different perspectives and measure how well models follow instructions and display interactions.
Open → 2610.02205v1

Robot skills improve by practicing tasks found in past data

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Abstract: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/

Thu 1 OctRoboticsArtificial Intelligence
The gist
Building robots that can reliably do many different tasks usually requires a lot of human work to teach and control them. The authors created a method where a robot can improve its skills on its own by practicing tasks it identifies from previous experiences in simulation. Their method helps the robot learn new helpful skills and fix mistakes without changing its main program. This process makes robots much better at completing tasks, even in the real world, after practicing many times in simulation.
Open → 2610.02204v1

Embedding prediction improves image generation quality in diffusion transformers

Embedding Prediction Helps Image Generation

Abstract: In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Thu 1 OctComputer Vision and Pattern RecognitionMachine Learning
The gist
Generating images using artificial intelligence often relies on giving the AI a starting instruction like a word or category. Usually, this instruction stays the same as the AI progressively refines the image. The authors show that predicting these instructions anew at each step, based on how the image looks at that point, helps improve image quality. They build a model that predicts multiple instruction points simultaneously and then used these predictions to guide image creation dynamically. Their approach achieves better results on a standard image dataset while using less training effort compared to some previous methods.
Open → 2610.02203v1

Benchmark predicts papers that spark new research ideas

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Abstract: What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

Thu 1 OctArtificial IntelligenceComputation and LanguageInformation Retrieval
The gist
Scientists are very good at finding old research papers that help solve new problems, a skill that AI has yet to match. To study this ability, the authors created ScholarCatalyst, a collection where lead authors marked which earlier papers helped or could have helped with their recent work, including reasons why. They tested computer programs on how well they could find these helpful papers using only information available at the time the new research began. Even advanced AI systems struggled to do better than simple search methods, pointing to the need for better training. This work aims to help build tools that can guide researchers to the right past studies when they have a new idea.
Open → 2610.02202v1

Silsa improves 3d shape generation with fewer tokens and better structure

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Abstract: High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

Thu 1 OctComputer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
The gist
Generating detailed 3D shapes usually involves breaking them into many tiny pieces, which can be slow and can mess up delicate connections. The authors created SILSA, a new way to represent 3D shapes using overlapping slices that keep surfaces continuous and connected. SILSA uses fewer tokens, making generation faster and more memory efficient, while better preserving thin parts and overall shape connections. Their tests show SILSA achieves higher accuracy and better shape quality compared to previous methods.
Open → 2610.02201v1

Vista harness improves multimodal reasoning in visual environments

VISTA: A Visual Harness for Reasoning in an Interactive World

Abstract: We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

Thu 1 OctArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Some AI models can understand and reason about pictures, but they struggle with long sequences of events. The authors created VISTA, a tool that helps these models remember and organize what they see over time, allowing better decision-making. With VISTA, an AI model performed perfectly on a set of visual games and used fewer actions than humans did. This tool can also work well across different types of visual puzzles and games with little adjustment.
Open → 2610.02200v1

TACO optimizer slashes memory use for large language model tuning

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Abstract: Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized $1\to1$ operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by $174\times$ relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by $2.9\times$ (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable accuracy and runtime. TACO further enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 GPU across multiple model families and tasks.

Thu 1 OctMachine Learning
The gist
Fine-tuning large language models usually requires a lot of memory for storing optimizer information, which limits how big models can be on current GPUs. The authors propose TACO, a new optimizer that drastically reduces this memory need by storing only a small, simple piece of gradient information per column of model weights. This approach keeps accuracy and speed similar to current methods while allowing much larger models to be fine-tuned on a single GPU. TACO lets teams work with bigger language models without needing more powerful hardware upgrades.
Open → 2610.02199v1

Policy optimization method avoids unstable action gradients in reinforcement learning

FERPO: Forward Entropy-Regularized Policy Optimization

Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).

Thu 1 OctMachine LearningArtificial IntelligenceRobotics
The gist
When teaching a computer to make decisions in complex situations, one way is to guess how good each action is and improve based on that guess. But predicting the exact value of an action doesn't always help improve decisions correctly because the details can be unreliable. The authors propose a new method called FERPO, which improves decision-making without needing those tricky details by comparing current choices to a carefully balanced target. This approach encourages exploring multiple good options and is more stable and efficient in tests on simulated robotic control tasks.
Open → 2610.02198v1

HiPhy improves video generation with realistic physical interactions

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Abstract: Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Generating videos that follow real-world physics is hard, especially when multiple physical effects happen together. The authors propose HiPhy, a new method that makes videos where different physical rules work correctly at the same time, like a balloon floating while steam rises. They created a large dataset and benchmark to test this idea. Their method does better than previous ones at making videos that look physically and semantically believable.
Open → 2610.02197v1

Humanoid robot controllers evolve skills at test time for new tasks

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Abstract: We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

Thu 1 OctRoboticsComputer Vision and Pattern RecognitionGraphics
The gist
Robots can struggle to handle tasks they weren’t specifically trained for. The authors show a way for humanoid robots to adapt on the fly by tweaking how they try to reach goals without retraining their brains. They combine a special model that understands the robot and objects with reward programs that can be changed by an AI assistant guided by feedback from the robot’s actions. This lets the robot discover new ways to use what it already knows to solve different tasks, both in simulations and on a real robot.
Open → 2610.02196v1

Exactly solving cost-aware mass transport on graphs with matrix methods

Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control

Abstract: The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, with a temporal-difference penalty that restores the cost. A state cost folds into the reference process as a Feynman-Kac tilt. The cost-augmented bridge is then a plain bridge against the tilted reference, and the penalty is unnecessary. The bridge is computed exactly by alternating two endpoint rescalings, each one sparse matrix-exponential application; nothing is discretized in time or learned. The alternation converges at a rate set by the endpoint coupling alone. For a quadratic congestion cost on time-averaged occupancies, damped best response around the exact bridge is gradient descent on a strongly convex function, and its residual bounds its error. On a protein-folding model, a free-energy cost lowers the expected barrier of the folding paths. On the learned approach's road network, roll-outs of the exact bridge match the target within sampling error, and on networks with millions of intersections its memory grows linearly.

Thu 1 OctMachine Learning
The gist
Moving resources or things between points on a network often has costs depending on the path taken. The authors showed a new way to find the best paths exactly without learning or approximations by changing the underlying network with a special mathematical tilt. This approach works faster and with less memory, even on huge networks, and can model complicated costs like congestion. Their method matches or improves on previous learned approaches and gives clear error bounds during optimization.
Open → 2610.02195v1

Hierarchical continuous diffusion models improve puzzle and language tasks

Hierarchical Continuous Diffusion Language Models

Abstract: Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.

Thu 1 OctComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Generating sentences or solving puzzles with computers is tricky because words depend on each other. The authors found a way to combine two methods: discrete words and flowing hidden signals, so the computer can consider the whole sentence at once while creating it step by step. This method improves accuracy in solving puzzles like Sudoku and planning problems, as well as making better guesses for the next words in text. Their approach keeps the hidden signals as the main state, reading and updating words from it continuously. This helps the computer keep words related and consistent while generating text or solving problems.
Open → 2610.02193v1

Large language models improved by fixing math reasoning steps

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.

Thu 1 OctMachine Learning
The gist
Solving math problems with AI is tricky because it needs clear step-by-step understanding, not just guesses. The authors studied what kinds of math reasoning large language models (LLMs) can and cannot do well, finding that discovering new math ideas is the hardest part. They created a test to measure different reasoning skills and found ways to fix common mistakes after training. Their new method helps models learn better math reasoning by focusing on these key reasoning parts, improving their problem-solving skills.
Open → 2610.02191v1

Adaptive step size improves large model training stability and speed

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Abstract: Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.

Thu 1 OctMachine Learning
The gist
Choosing how big a step to take when training large AI models is tricky: too small means slow learning, too big can cause problems. The authors combine two methods to pick the direction using known gradient info and then decide the step size by checking the model’s performance nearby. This helps adjust steps without expensive calculations and keeps training stable and efficient. Their method works well across different models and data compared to always using fixed step sizes.
Open → 2610.02190v1

Idiom model and rl-sae method generate and control disordered protein sequences

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Abstract: Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.

Thu 1 OctMachine Learning
The gist
Intrinsically disordered protein regions (IDRs) are flexible parts of proteins important for many cell functions but are hard to design because they don't have fixed structures. The authors created IDiom, a language model trained specifically on these disordered regions, to generate realistic sequences. They also developed a method called RL-SAE that helps control specific sequence features linked to function by rewarding the model when it generates desired patterns. This approach produces sequences that better match desired biological activities than previous methods. Their tools allow more interpretable and targeted design of these flexible protein regions.
Open → 2610.02189v1

Fast image and video generation with fewer steps using adversarial distillation

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Abstract: Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.

Thu 1 OctComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Generating images or videos from AI models usually takes many steps and a lot of computing power. The authors found a way to train a simpler AI model to generate visuals in just a few steps by teaching it to match distributions using adversarial methods. Their approach avoids extra costly calculations and improves the quality of the generated images and videos significantly. Their method works well on standard benchmarks and beats previous few-step techniques.
Open → 2610.02188v1

Higher-order grammar improves molecule generation and learning quality

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Abstract: Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.

Thu 1 OctMachine LearningArtificial Intelligence
The gist
Molecules have complex shapes that are hard for computers to understand fully, especially parts like rings or common patterns. The authors created a new way called Higher-order Grammar Representation that breaks molecules down into simple rules capturing these complex parts. This method works well with existing sequence-based computer models, making molecule generation and learning more accurate and efficient. They also made a new big ring-focused set of molecules to better test such methods. Their approach leads to perfect molecule validity in generation and better results in predicting molecular properties.
Open → 2610.02186v1

Looped transformers improve decoding efficiency without extra training

Decoding Looped Transformers Better for (Almost) Free

Abstract: Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.

Thu 1 OctMachine Learning
The gist
Looped transformers use the same part of a model multiple times to save parameters. Usually, only the last pass helps pick the next word, but earlier passes contain useful clues that are wasted. The authors introduce LoopCD, a method that compares early and late passes during decoding to make better choices without extra training. This improves performance while allowing fewer passes, which saves computing work.
Open → 2610.02185v1

Unified explanations clarify cause and effect in reactive systems

Sufficient Reasons and Explanations for Reactive Systems

Abstract: We address the problem of temporal causality and explainability for reactive systems, and, in this setting, study sufficient reasons and contrastive explanations. These two notions are well-known explainability measures in the context of neural networks. In this work, we unify these notions for reactive systems and formal specifications given in temporal logic, providing dedicated definitions for sufficient reasons and contrastive explanations. We then lift these definitions to \emph{temporal} sufficient reasons and contrastive explanations, providing more general and symbolic representations of explainability. We analyze the complexity of both verifying and finding explanations of the different types, and we demonstrate our approach using a prototype implementation.

Thu 1 OctFormal Languages and Automata Theory
The gist
Understanding why a system does something over time can be tricky. This paper looks at how to explain the behavior of systems that react continuously to inputs, like software controlling machines. The authors combine two ways of explaining decisions, called sufficient reasons and contrastive explanations, and make them work for systems described using time-based rules. They also analyze how hard it is to find and check these explanations and show some examples using their own tool.
Open → 2610.02184v1

SoftServe improves deep learning training with scalable quasi-Newton methods

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Abstract: Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.

Thu 1 OctMachine LearningArtificial Intelligence
The gist
Optimizing deep neural networks can be very hard because they have many parameters and their math is complicated. The authors created SoftServe, a new way to improve training by estimating curvature to guide updates better, even when the usual assumptions don’t hold. It uses efficient math tricks that work well with GPUs and can handle very large networks. This approach often achieves better performance than popular existing methods on difficult learning tasks.
Open → 2610.02182v1

OmniSeek improves multi-turn audio visual reasoning with built in tool use

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

Abstract: We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Understanding important details in long videos or audio clips can be hard because there is too much information. The authors created OmniSeek, a smart system that actively chooses when and where to look or listen to find key clues before making decisions. It learns by practicing on lots of example reasoning steps involving both sounds and images and gets better through a special kind of trial-and-error training. OmniSeek can combine audio and visual clues over several turns to make better judgments than if it looked at everything at once.
Open → 2610.02181v1

Generative cinematographer lets artists control 3d camera and object motion

Generative Cinematographer: Composing Camera and Object Motion in 3D

Abstract: Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.

Thu 1 OctComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
It is hard to control video scenes because moving a camera and objects in two dimensions can mean many different things in three dimensions. The authors created Generative Cinematographer (GenCine), a system that turns a single image into a 3D scene where artists can move the camera and parts of objects in 3D space. This system uses colored handles to let artists control how objects move piece by piece, even for complex shapes without needing special physics tools. GenCine then uses a special method to turn these 3D controls into instructions for a video generator, so the final videos follow the artist's motions realistically. It works on real and synthetic videos and keeps object shapes consistent when the camera changes viewpoint.
Open → 2610.02180v1

Multi teacher distillation reveals factors shaping reinforcement learning updates

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Abstract: Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.

Thu 1 OctMachine Learning
The gist
Combining the guidance from multiple AI teachers into one student model can be tricky, especially when these teachers learn by trial and error. The authors studied how different teacher signals change the student’s learning by looking closely at how the model’s parameters update in practice. They found that the way losses are averaged, the behavior of the optimizer, and number formats all affect how the student learns from its teachers. These insights help explain why some approaches to blending teacher advice work better than others in specific tasks like math.
Open → 2610.02179v1

Effective resistance predicts reliability in tissue protein networks

Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes

Abstract: Protein function annotation needs to know which predictions to distrust, not only what a model predicts. We ask whether tissue-specific interaction structure carries that information. Our candidate signal is effective resistance, used previously to relieve over-squashing by rewiring. Across 24 tissue-specific interactomes it is dominated by inverse degree, and the degeneration deepens as the co-expression filtered network grows, with a Spearman correlation of -0.955. The residual departure from that limit exceeds degree-preserving null graphs in all 24 networks. Controlling for predictive entropy, degree, annotation cardinality, local structure and feature-only difficulty, the residual explains additional per-node loss in 19 of 24 held-out networks once a permutation floor is subtracted, at every depth, and the effect strengthens monotonically with depth. The increment reaches 0.37% of the variance the controls leave unexplained, 5.6 times a permutation floor, against 1.5 times when the model is retrained in a degree-preserving null world. Selective prediction improves negligibly. The signal is reproducible; degree degeneration bounds it.

Thu 1 OctMachine Learning
The gist
Predicting how proteins work in different body tissues is hard, and it’s important to know when those predictions might be wrong. The authors studied special maps of protein interactions within tissues and found that a measure called effective resistance can help identify where predictions are less reliable. They discovered this measure mostly reflects simple properties of the network but still adds useful insight beyond those. This insight improves understanding of model reliability as predictions get more detailed.
Open → 2610.02175v1

Language model parts adjust predictably after removal of signals

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Abstract: Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

Thu 1 OctMachine LearningComputation and Language
The gist
When a piece of a language AI is removed, other pieces seem to adjust as if fixing the problem. The authors explain that this apparent self-fixing is actually a predictable response already present before removal. They found a simple mathematical rule that describes how other parts increase or decrease their activity when one part is changed. This rule holds true across several AI models and different components inside those models. So, the so-called self-repair is actually just normal balancing behavior in the model's parts.
Open → 2610.02173v1

Robots infer partner limits to work together on new tasks

Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination

Abstract: Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry alone. Physical limitations caused by hardware degradation or actuator faults can restrict the actions a robot can reliably execute, yet these limitations may be unknown to its partner. We study whether a helper can infer a robot partner's physical constraints from observing it coordinate with another robot, then use the inferred capability to coordinate with the same partner on a new task. This is difficult because a demonstration shows what the constrained robot did, but not what it could have done. In physically coupled tasks, the other robot may also compensate for its limitations, making those limitations difficult to identify from the constrained robot's behavior alone. Our key insight is that these constraints shape the joint behavior of the team, making the actions of both robots informative about the constrained partner's capability. We introduce Watch, Infer, Coordinate, a benchmark spanning three physically coupled manipulation settings, together with an inference approach that scores candidate constraints using observed joint behavior. Across all three settings, our method substantially improves constraint inference and zero-shot coordination, approaching an oracle with access to the true constraints.

Thu 1 OctRoboticsArtificial IntelligenceMultiagent Systems
The gist
Robots often need to team up to move big or heavy objects, but sometimes one robot has hidden physical problems that limit what it can do. The paper studies how a robot can guess its partner's limitations by watching how they work together with another robot. By understanding these limits, the helper robot can better coordinate with the partner on new tasks without needing prior knowledge. The authors introduced a way to measure these limits from their shared behavior and tested it in different scenarios, showing it works well.
Open → 2610.02170v1

Quantum impurity models can be efficiently simulated classically and quantumly

Polynomial-time classical and quantum simulation of quantum impurity models

Abstract: Quantum impurity models are paradigmatic models of interacting quantum matter, as well as key computational primitives for modern electronic-structure methods. They describe a small subsystem of interacting fermions coupled to a large, noninteracting bath. We perform a comprehensive study of the computational complexity of simulating impurity models, delineating the boundary between classical and quantum tractability for this class of problems. Our main finding is that static properties of quantum impurity models can be calculated efficiently on a classical computer. Specifically, we give classical algorithms that (1) estimate the ground-state energy to additive precision $δ$ in time $\mathrm{poly}(n,δ^{-1})$, and (2) estimate the partition function at inverse temperature $β$ to relative precision $δ$ in time $\mathrm{poly}(n,β,δ^{-1})$, where $n$ is the system size. These results improve the previous best-known complexity for ground-state energy estimation from quasipolynomial to polynomial time, while establishing for the first time rigorous polynomial-time guarantees for simulating impurity models in thermal equilibrium. On the other hand, we find that simulating dynamical properties of impurity models is hard for classical computers but easy on a quantum computer. As a canonical example, we show that computing their nonequilibrium Green's functions captures the full power of quantum computation, even at finite temperature. Taken together, our results rule out superpolynomial quantum speedups for computing static properties, but provide an avenue for quantum advantage in simulating impurity physics out of equilibrium.

Thu 1 OctComputational ComplexityData Structures and Algorithms
The gist
Quantum impurity models describe a small interacting system linked to a large environment. The authors show that key static properties of these models, like energy, can be calculated efficiently using regular computers. However, calculating how these systems change over time requires a quantum computer and cannot be done quickly on classical ones. This work clarifies what parts of simulating these models are easy or hard for classical and quantum computers.
Open → 2610.02167v1

Quantum spin glasses require complex circuits for state preparation

Beyond Light Cones: State Preparation Complexity in Quantum Spin Glasses

Abstract: We introduce a method for studying state preparation complexity in dense quantum $p$-spin Hamiltonians on $n$ qubits, going beyond bounds based only on circuit lightcones. The key input is the class's effective profile complexity, which is derived from the metric entropy of its Pauli profiles. These profiles record expectations of all Pauli operators supported on exactly $p$ qubits. Classes with uniformly bounded quadratic effective profile complexity remain separated from the ground-state energy by a positive multiple of $\sqrt n$ for sufficiently large fixed $p$. At subquadratic effective profile complexity, the class cannot outperform a suitable benchmark class at leading order, with product states providing a universal benchmark. The proof combines an adaptation of a nonsymmetric quantum de Finetti theorem of Berta et al. (arXiv:1810.12197) with Gaussian process entropy bounds. Applying this framework, we show that attaining near-ground-state energy requires $Ω(n^2/\log n)$ one- and two-qubit gates, even with arbitrary discardable ancillas. We also obtain depth-width tradeoffs, entanglement-depth and matrix product state bond-dimension lower bounds, and obstructions for both orientations at every fixed level of Parham's magic hierarchy (arXiv:2504.19966), with total circuit width $O(n)$. In first-level reverse magic, a shallow circuit is followed by an unrestricted Clifford circuit. The latter can spread local observables across the system, preventing a direct application of small-lightcone bounds. For this first-level class, our bounds also allow arbitrarily many clean ancillas at fixed shallow-circuit depth. A sharper benchmark shows that Clifford+$T$ circuits with $o(n)$ $T$-gates have no leading-order energy advantage over product stabilizer states, even with unrestricted Clifford operations and arbitrary discardable ancillas.

Thu 1 OctComputational Complexity
The gist
Preparing the lowest-energy states of dense quantum spin glasses is very hard and needs complicated quantum circuits. The authors show that simple or shallow circuits cannot get close to these states, even if you have extra helper qubits. They introduce new mathematical methods to prove that many quantum gates and a lot of circuit depth are necessary. This helps us understand fundamental limits on how quantum computers can simulate such complex systems.
Open → 2610.02166v1

Coding agent learns when and how to compact context for better long tasks

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Abstract: Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.

Thu 1 OctComputation and Language
The gist
Coding agents work on software tasks that need many steps, but remembering too much can be a problem as they go. The authors created AutoCompact, which helps these agents decide when to shrink their memory and what details to keep for later. They trained the agent by correcting poor memory choices and then improved it with rewards when the overall task succeeded. This method led to better performance in coding tests, even when the agent's memory was limited or very large.
Open → 2610.02163v1

World observer improves tracking of objects outside agents view

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

Abstract: How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor's view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Video world models usually focus only on what an agent currently sees, losing track of objects that move out of view. The authors present World Observer, a system that lets the model watch parts of the environment independently from the agent’s viewpoint. This means objects outside the agent’s view keep updating correctly and look consistent when seen again. Their method uses multiple perspectives anchored to a shared panoramic view to keep track of changes anywhere in the scene. It performs better at predicting how things change out of sight while keeping good visual quality and adherence to 3D space.
Open → 2610.02162v1

Multi-robot coordination improved by semantic communication framework

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.

Thu 1 OctRoboticsArtificial Intelligence
The gist
Coordinating multiple robots to work together on complex tasks is difficult because they need to plan carefully and act precisely. The authors introduce DuoMind, a system where each robot thinks about the big picture using language and vision models and manages actions and messages with other robots. This approach helps the robots coordinate better over long tasks. The authors also created RoboPoly, a set of tests to measure how well robots work together on these tasks. Their experiments show that DuoMind helps robots complete multi-robot jobs more effectively.
Open → 2610.02161v1

4Director improves control of video scenes using 3D object geometry

4Director: Controlling Video World Models with Rigid 3D Geometry

Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Thu 1 OctComputer Vision and Pattern Recognition
The gist
Controlling how cameras and objects move in videos is tricky because existing methods either only work roughly or lose important 3D details. The authors created 4Director, which builds a detailed 3D model of each object from an input image and moves it precisely using fixed 3D transformations. This way, the video remains consistent from different viewpoints and the shapes don’t change unexpectedly. They also made a special system to turn this 3D info into high-quality videos and a new dataset to train and test their approach. Their experiments show their method controls objects and cameras better than earlier techniques.
Open → 2610.02160v1