Week beginning 28th September 2026

Every computer science paper posted to arXiv this week, with plain-language summaries and practical uses for each one. Includes commercial applications where relevant.

Model segments entire 3d shapes into parts from point guides

Point2Part: Unified 3D Partitioning from Point Prompts

Abstract: Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Dividing 3D shapes into parts is tricky because parts can overlap or leave gaps. The authors present a method that splits a whole shape into non-overlapping parts based on points users pick on the shape. Their model works with images or meshes and treats all parts at once to ensure they perfectly cover the shape without overlap. This approach leads to better, more reliable part segmentation than previous methods.
Open → 2609.38180v1

Robots improve their actions using reusable short skills autonomously

Skill-Space Shooting for Autonomous Robot Policy Improvement

Abstract: Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.

Tue 29 SeptRoboticsArtificial IntelligenceMachine Learning
The gist
Robots often face new problems they weren’t trained for, so they need to learn how to fix mistakes on their own. The authors propose a way for robots to use a set of simple, familiar actions, called skills, to try different fixes guided by AI models. When a skill works well, the robot learns from it to improve how it handles the task next time. This method helps robots get better at many tasks without needing humans to show every correction. Experiments show robots can keep improving and share skills to learn faster on new tasks.
Open → 2609.38178v1

Multimodal language models imagine 3d scenes for better spatial reasoning

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

Tue 29 SeptComputer Vision and Pattern RecognitionComputation and Language
The gist
Understanding 3D scenes from different pictures is hard for AI that reads words and images together. The authors noticed that people don’t think about every tiny detail but instead imagine a simple 3D sketch of objects and their positions. They taught a language-and-image AI to do something similar by creating a simple 3D summary before answering questions. This method helped the AI understand space and objects better than before.
Open → 2609.38177v1

Semantic information links timing of language model decisions

Breakdown of Local Denoising as Semantic Speciation

Abstract: The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a "common cause" hypothesis, we prove that the nonlocality window must lie in the speciation window. This hypothesis postulates that semantic labels explain a fraction of the correlations between distant tokens, a condition that is natural for many real datasets. We further give conditions under which both windows shrink to a single limiting time as system size grows, defining a "phase transition", and verify this behavior analytically in Gaussian mixtures. Together, these results identify conditions under which semantic information explains the concurrence of speciation and nonlocality, connecting two complementary perspectives on the emergence of semantic structure in generative modeling.

Tue 29 SeptMachine Learning
The gist
Generative models that create text or images make decisions in stages: first, they decide what kind of thing they're generating (like choosing a category), then they fill in the details. This paper studies when exactly these two stages happen and finds that the stage where the model needs information from far away in the input overlaps with the stage where it decides the category. The authors prove this overlap under certain natural conditions and show that as models get larger, these stages merge into one sudden change. This helps explain how models organize meaning during generation.
Open → 2609.38176v1

Robot learns manipulation tasks from simple visual examples

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Abstract: We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.

Tue 29 SeptRobotics
The gist
Robots often need many examples and complex training to learn tasks. This work clearly defines what robots should learn from watching visual demonstrations and offers a simple system called SimpleICL that doesn't require expensive data or pretraining. The system learns to recognize actions, objects, and how to use them together, working well in both simulations and real-world tests. The authors will share all their data and tools so others can build on their work.
Open → 2609.38173v1

Humanoid robots learn to carry diverse objects from few videos

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Abstract: Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

Tue 29 SeptRoboticsComputer Vision and Pattern RecognitionGraphics
The gist
Teaching humanoid robots to carry and manipulate objects by watching videos is hard because it's difficult to get many good examples. The authors developed a method called PRISM that takes just a few real videos and creates many new, varied ones by changing object interactions. These videos are then turned into realistic motions, allowing the robot to learn how to carry different objects it hasn't seen before. The robot successfully picks up and carries items like boxes and balls without extra real-world training.
Open → 2609.38172v1

Adversarial training improves fine details in pixel image generation

Adversarial Training for Pixel Diffusion

Abstract: Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Pixel diffusion models can create images by predicting pixels directly but often miss fine details seen in real photos. The authors show that adding adversarial training after initial training helps these models produce sharper and more natural textures without changing how they generate images. This method fixes missing high-frequency details important for realism, unlike other approaches that add noise or lose alignment with the image description. The improvement works well when the model directly outputs the image pixels, but not when using compressed forms.
Open → 2609.38170v1

StepQuant reduces memory use for AI with smarter state compression

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Abstract: Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

Tue 29 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Modern AI language models remember information during output generation using states that can take up a lot of memory. The authors found that errors in compressing these states affect the model’s accuracy differently depending on when and where they happen. They created STEPQuant, a method that decides how much detail to keep in each part of these states to reduce memory without losing accuracy. Tests showed STEPQuant keeps the same quality as the original but uses much less memory, making AI models easier to run efficiently.
Open → 2609.38169v1

LeapQuant speeds up linear attention for large language models

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Abstract: Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.

Tue 29 SeptMachine LearningArtificial Intelligence
The gist
Large language models use a process called attention to understand context, but it can be slow and computationally heavy. Some newer designs use a smaller fixed-size memory to handle long texts more efficiently, but updating this memory repeatedly still takes time. The authors propose LeapQuant, a method that compresses this memory using fewer bits without losing accuracy, by updating it less often and carefully handling data that could cause errors. This allows the models to work faster and use less memory during inference while maintaining quality.
Open → 2609.38166v1

Transformer model improves crop type mapping from satellite images

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Abstract: The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.

Tue 29 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Knowing exactly which crops grow where helps farmers and food planners. The authors created a new computer model called PAtteRNS that looks at satellite images in three ways—time, color, and space—to better identify different crop types. Their method is faster and more accurate than others, especially at detecting crop field edges. They also found that how datasets are prepared affects results a lot and suggest better standards are needed for future improvements.
Open → 2609.38165v1

Rho models improve robot hand coordination and task adaptation

Rho: A Foundation for Efficiently Adaptable VLA Models

Abstract: General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.

Tue 29 SeptRobotics
The gist
Making robots with two arms work well on many tasks is hard because each robot looks and moves differently. The authors created Rho, a set of robot control models that work with several popular two-arm robots. These models learn well even with little new data and can quickly adjust to new tasks. They tested Rho both in simulations and real robots, showing it matches or beats existing systems. Rho can also adapt on the fly by learning from human corrections, which helps it handle tricky situations it didn’t see before.
Open → 2609.38164v1

World action models improve robot learning with calibrated representations

Rethinking Representations for World-Action Modeling

Abstract: World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

Tue 29 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Teaching robots to act and predict what they will see next is tricky because it depends on how the robot understands the world. The authors found that just making the robot’s visual predictions very accurate or using pre-made visual features doesn't guarantee good actions. They created a new method called ReWAM that organizes these visual features in a way that helps the robot learn better policies by letting the robot’s actions shape what it pays attention to. This approach improved robot task success on two robot training setups without needing extra video training.
Open → 2609.38163v1

Scaling regret matching converges fast to equilibrium in zero sum games

Why Last-Iterate Scale-Invariant Regret Matching Converges Linearly?

Abstract: IREG-PRM+ normalizes the cumulative regret vector by its own norm and attains optimal regret without knowledge of the payoff scale. Run unmodified on zero-sum matrix games, it converges linearly in the last iterate, and no analysis explains why. The obstacle is that the algorithm has no fixed step size to analyze: the step size is a state variable, the inverse of a regret norm that the trajectory itself moves. Every proved linear rate for regret-matching dynamics comes from restarting or modifying the update. We identify the mechanism as norm saturation: the regret norm rises to a finite limit and freezes the step size. We prove that it always does, with an explicit bound, and that saturation forces the last-iterate Nash gap to vanish on every matrix game; pointwise convergence follows whenever the equilibrium is unique. Near a unique strictly complementary equilibrium the active support freezes in one step, and the one-round Jacobian on that support has a closed form. The last-iterate then converges linearly at a closed-form rate, provided one scale-invariant quantity stays below one: the saturated step size times the largest singular value of the value-centered payoff submatrix on the support. On the $216$-instance testbed, the $184$ instances with a resolvable limit all satisfy it. The same analysis gives a ratio certificate: observable norm-increment ratios bound the unobservable Nash-gap ratio up to a constant that enters once and does not accumulate with the iteration count. Its slope-two law holds on $96.1\%$ of the instances where the slope is measurable, and the same increment monitors progress in extensive-form games, where best-response passes can be scheduled sparsely. The code is available at https://github.com/lbn187/NormCert.

Tue 29 SeptComputer Science and Game Theory
The gist
This paper explains why a particular algorithm called IREG-PRM+ quickly finds the best strategy in zero-sum games without knowing how large the payoffs are. The authors show that the algorithm’s step size naturally stabilizes because the cumulative regret it tracks hits a limit, causing the process to lock in a steady pace. They prove this stabilization always happens and that it guarantees the algorithm’s decisions get steadily closer to the game’s perfect balance point. The paper also provides a way to monitor progress by observing certain measurable ratios, making it easier to track performance even in more complex games.
Open → 2609.38162v1

Spectral distance bounds reveal patterns in reconstructed graphs

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Abstract: Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends coincide exactly when the reconstruction only adds edges or only deletes them, and on that class the distance is a rescaled edge count that says nothing about which edges changed. When the ends differ, the residual between the distance and the lower end is positive only if the reconstruction both invented and lost edges, which turns it into a certificate of mixed editing computable from the reported summaries alone. We characterize these regimes in 135 reconstructions produced by three open-weight models over 45 synthetic graphs. Seventy-seven outputs are one-sided and 29 mixed outputs have $X > 0$, including cases where edge count is exactly preserved while nineteen edges were simultaneously invented and lost. The three models differ in editing policy, ranging from copying the input to attempting completion at the cost of large hallucination volume, a distinction that aggregate distortion does not reveal.

Tue 29 SeptMachine LearningDiscrete Mathematics
The gist
When computer programs try to recreate networks, it’s hard to know how close the new network is to the original. The authors found a way to measure this by looking at certain numbers related to the network’s structure, which can tell if edges were only added, only removed, or both. They tested these ideas on many examples, showing different models behave differently when rebuilding networks. Their work helps us better understand and compare how models change networks during reconstruction.
Open → 2609.38161v1

Information splits explain probability deviation events geometry

An information identity reveals the geometry of deviation events

Abstract: In a deviation event, the empirical measure of $n$ independent draws lands in a set of distributions. The exponent of its probability splits, at every $n$, into two nonnegative terms. The first is the rate: $n$ times the relative entropy, from the population, of the distribution of a typical draw under the event. The second is the dependence that conditioning induces among the draws, measured by their total correlation. Their relative sizes depend on the geometry of the set. This paper focuses on this exact split and on the geometry that the interplay of its two terms reveals. The dependence vanishes exactly when the event confines every draw to a single set, and the classical rate of that constraint is exact. On a set invariant under a compact group that fixes the population and leaves no invariant set of intermediate probability, the rate vanishes; uniform convergence over a hypothesis class that such a group permutes is one. On a half-space whose threshold sits a fixed number of standard errors above the mean, the fraction of the exponent that is dependence tends to a function of the event's probability alone. On a set split into pieces, the dependence is the pieces' average plus $n$ times the Jensen-Shannon divergence of their marginals, minus the entropy of their weights. The split extends to a general reference law. When the information projection is such a law, the split gives the exact value of the expectation in the exponential change of measure to the projection.

Tue 29 SeptInformation Theory
The gist
When you randomly draw samples, sometimes the results deviate from what you expect. This paper finds that the chance of such deviations can be exactly split into two parts: one measuring how different the most likely outcome is from the usual, and another measuring how the samples become dependent when you know the event happened. The balance between these parts depends on the shape of the event you consider. The authors also explore how this split works under different conditions and sets, revealing hidden structure in how deviations occur.
Open → 2609.38159v1

EmoRES improves emotional speech control in text to speech systems

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Abstract: Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.

Tue 29 SeptSoundComputation and Language
The gist
Sometimes, computer systems that turn written text into spoken words struggle to express emotions correctly. The authors found that an emotion signal inside these systems can be broken into two parts: one that just changes speech away from neutrality, and another that steers it toward a specific emotion. They created a method called EmoRES that adjusts these parts separately without needing extra training. This approach helped machines produce speech that better matches emotions and sounds more natural, according to tests and listener feedback.
Open → 2609.38157v1

Pixel space training improves few step text to image generation

DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses

Abstract: Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Generating images from text usually takes many steps and complex systems working on compressed image data. The authors explore a way to speed this up by training models directly on normal RGB images (the pixels you see) instead of compressed forms. They find that matching textures at some noise level and using a new way to guide learning with existing visual representations helps create good images faster. Their new method, called DMA², produces high-quality images in fewer steps than previous models, making fast text-to-image creation more practical.
Open → 2609.38156v1

Grounded biographies improve question answering about long videos

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Abstract: Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

Tue 29 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
The gist
Answering questions about long videos is hard because the same object can appear many times and look similar to others. The authors created a method that links all appearances of a single object across a video into a ‘biography’ so the system can track it over time. This helps answer questions about events involving that object more accurately. Their approach worked better than previous methods on tests with day- and week-long videos.
Open → 2609.38155v1

Video generation models gain reusable tools for faster diverse tasks

LongLive-Plug: Once-for-All Distillation for Video Generation

Abstract: Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Video generation models often need extra training when adapted for different jobs, which takes time. The authors created LongLive-Plug, a way to train helpful add-ons just once on a main model. These add-ons can then be plugged into many different video models without retraining, making tasks like faster sampling and longer video correction easier. This approach was tested on many models and tasks, showing the add-ons work well across different settings.
Open → 2609.38154v1

Power diagram based simulation and rendering for dynamic 3D scenes

PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams

Abstract: We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural alignment between the two: the geometric and appearance properties of each primitive correspond closely to the quantities MPM already tracks as an object deforms. Consequently, simulated motion can drive the scene's geometry and appearance directly, without an auxiliary representation in between. Built on this framework, we enable a range of applications on real and synthetic scenes: (1) simulating a static scene under user interaction, (2) recovering spatially varying material fields, (3) compositing primitives from independently captured scenes into a single simulation-ready scene and (4) ray-tracing reflections that update consistently as the object deforms. Our results suggest that PowerSim excels over previous frameworks for physically grounded dynamics, while unlocking unique advantages-such as secondary ray lighting effects on dynamic scenes. Results are best viewed on our project website: https://power-sim.github.io/.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Simulating realistic motion and appearance changes in 3D scenes is difficult because geometry and physics must work together. The authors present PowerSim, which links a special 3D representation called PowerFoam with a physics simulation method called Material Point Method. This connection lets the simulated movements directly update the scene’s shape and look without conversions. PowerSim also allows interactive scene changes, material property recovery, combining different scenes, and realistic lighting effects that adjust with motion.
Open → 2609.38153v1

FracGen generates realistic crack and tear videos from single images

FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation

Abstract: We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to capture physical state rather than surface appearance; and it is supervised with physics-informed losses that encourage consistency among the predicted maps. As a result, FracGen captures distinct material-specific fracture behavior without expensive test-time simulation or per-scene tuning, while offering fine-grained control over where an object tears, how fast the crack propagates, and how much deformation precedes failure. We further introduce a benchmark for evaluating the physical plausibility of generated fracture video, and show through extensive experiments that FracGen outperforms existing video generation baselines in both physical and visual fidelity. Results are best viewed in our project website: https://fracgen.github.io/.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Predicting how objects break or tear when stressed is hard, especially from just one picture. The authors created FracGen, a tool that learns to generate believable videos of objects cracking and stretching by training on simulated fracture data. This simulation, called FracSim, provides detailed maps about how damage spreads inside materials, helping FracGen learn the physics behind breaking. As a result, FracGen can produce videos showing realistic fractures that can be controlled by specifying how and where the object breaks, without needing expensive calculations each time.
Open → 2609.38152v1

Latent information feedback improves transformer language models performance

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.

Tue 29 SeptComputation and Language
The gist
Transformer language models usually process information in a one-way, feed-forward manner, which limits their ability to remember and use previous details efficiently. The authors propose a new model called LIFT that allows information from deeper layers to be reused in earlier layers during training. They do this by teaching the model to predict both the next word and extra helpful information about the upcoming words, using guidance from an existing language model. Tests show that LIFT models perform better on language understanding and reasoning tasks while also being more efficient during use.
Open → 2609.38149v1

Agentic meta reasoning improves long task control in large ai agents

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

Tue 29 SeptArtificial Intelligence
The gist
Large AI agents solving complicated problems need to decide how to work step-by-step, like picking what to build on next or when to stop. The authors introduce agentic meta-reasoning, which adds a smart controller that manages these decisions instead of just following fixed steps. This controller summarizes progress, explores options, evaluates their value, and chooses the best next move without rereading everything done before. Their tests show this approach helps AI perform better on complex programming, reasoning, and proof tasks by reusing earlier work and making better choices.
Open → 2609.38147v1

Video generation improved by layout control for large viewpoint changes

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Abstract: We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Generating videos from a single image while controlling how the camera moves is hard, especially when the camera shows parts of the scene that weren’t visible before. The authors developed a method called LIFT that lets users specify not just the camera movement but also what objects should appear and where in a future view, particularly the last video frame. To train this, they created a special technique called on-policy self-distillation to teach the system to follow sparse layout instructions. They also built a dataset to test large viewpoint changes and found that LIFT produces better videos with more precise content control.
Open → 2609.38146v1

Quantum channels show exact secrecy limits for private communication

The exact strong converse exponent for private communication over quantum channels

Abstract: We determine the exact strong converse exponent for secret-key transmission and generation over general finite-dimensional quantum wiretap channels, measuring reliability and secrecy jointly by squared fidelity to an ideal secret key. We introduce a novel Renyi private information whose regularization characterizes this exponent. An exact integral representation in terms of ordinary private information gives a uniform continuity bound, showing that the regularized quantity converges to private capacity as the Renyi order approaches one. This establishes for any wiretap channel an exponential fidelity decay at every rate above private capacity.

Tue 29 SeptInformation Theory
The gist
Sending secret messages through quantum channels can be tricky because eavesdroppers might listen in. This paper finds the exact rate at which the chance of keeping messages secret drops sharply if you try to send information faster than the channel’s private capacity. The authors introduce a new way to measure privacy that helps them describe this drop precisely. Their findings prove that if you exceed the private capacity, the reliability of secret communication fails at an exponential pace.
Open → 2609.38144v1

Builder agents learn reusable skills to improve AI task environments

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Abstract: Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

Tue 29 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
AI agents work better when the place they operate in supports them well. This paper looks at how one AI, called the Builder, can learn to create better 'workspaces' or environments for another AI, called the Target, without changing their internal programs. The Builder learns patterns called Meta-Skills, which help decide when the Target needs help and what kind of help to give. Using these Meta-Skills, the Builder builds supports that help the Target perform better on new problems. The authors find that these learned Meta-Skills improve performance significantly over not providing support or just giving static resources.
Open → 2609.38143v1

Advisors learn better advice for language models using self-distillation

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Abstract: A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.

Tue 29 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
Large language models can be helped by smaller advisor models that suggest what the big model should do next. This paper shows a way for the advisor to learn better advice over time by selectively learning from feedback about past advice, ignoring less useful corrections. The authors created a method called Advisor Self-Distillation (AdviSD) that improves the advisor by focusing on advice that really changes the big model’s behavior. Tests with popular language model advisors showed AdviSD gives more accurate and adaptable advice than previous methods.
Open → 2609.38142v1

Video diffusion model improves with split-role mixture of experts

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Abstract: Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

Tue 29 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Video generation models often struggle because they treat all parts of a video the same way, which doesn’t fit video data well. The authors show that existing methods force the model to use all parts evenly, which can break the natural grouping of video information. They propose a new model called SplitMoE that divides experts into two types: one focusing on big-picture meanings and the other on detailed visuals. This helps the model learn faster and create better videos by keeping related information together.
Open → 2609.38140v1

Language model harnesses tested for accuracy and speed on long contexts

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Abstract: Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68\% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.

Tue 29 SeptComputation and Language
The gist
Language models can use special methods called harnesses to handle very long pieces of text better, but current tests don’t show clear differences between these methods. The authors created a new set of challenges where models must search and think strategically over large amounts of information to answer questions accurately and efficiently. They found that even strong combinations of models and methods struggle with these tasks, and how efficiently a model works depends a lot on the specific method used. This shows that testing both accuracy and speed is important when evaluating how well language models handle long texts.
Open → 2609.38137v1

Style transfer method reduces content leakage while keeping style intact

CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

Abstract: Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Style transfer tries to make a picture look like it's painted in the style of another image, but often bits of the original picture sneak through, making the final image look messy. The authors found that pushing too hard to remove these bits often ruins the style, and trying to keep the style strong brings back unwanted content. They created a new method called CLeaR that balances this by carefully separating content and style using multiple image analysis tools without extra training. This approach results in better style matching and less leakage in the generated images.
Open → 2609.38136v1

Multi-agent flow matching generates objects meeting hard constraints

Multi-Agent Flow Matching with Decoupled Generative Guidance

Abstract: Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without relying on the simultaneously computed guidance inputs of other agents. To this end, we introduce DeGG-Flow, a general framework for multi-agent flow matching with decoupled generative guidance. By representing the generative process as a control-affine dynamical system, we develop guidance conditions for two classes of coupled requirements: shared requirements whose satisfaction depends on multiple agents together, and private requirements associated with each individual agent dependent on its neighbors. For both classes, we establish feasibility conditions and finite-horizon convergence guarantees. We further derive a Wasserstein bound that characterizes the distributional deviation induced by the guidance. We demonstrate DeGG-Flow on multi-robot collaboration for crossing a spatial gap by reconfiguring the environment, and on multi-object scene generation with affordance requirements. Across both applications, DeGG-Flow directly generates objects that satisfy all corresponding hard requirements, including at team sizes unseen during training.

Tue 29 SeptMachine LearningMultiagent SystemsRobotics
The gist
Creating multiple things at once often means they need to work together and follow strict rules, which is tricky. The authors introduce a method called DeGG-Flow that helps each part decide how to act without depending on others simultaneously. Their approach can handle both rules that involve the whole group and rules that only concern a few nearby parts. They show that using this method leads to generating objects or actions that always follow those strict rules, even when there are more parts than before.
Open → 2609.38133v1

Achieving constant rate optimality gap in budget constrained decision systems

Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

Abstract: We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$ optimality gap under general conditions, and has further identified conditions under which policies can achieve a better-than-$1/\sqrt{N}$ optimality gap. However, for general WCMDPs, no prior result achieves an optimality gap better than $1/\sqrt{N}$. In this paper, we identify conditions analogous to those for RBs under which a better-than-$1/\sqrt{N}$ optimality gap is achievable, and design a policy that attains an $O(1/N)$ optimality gap. Notably, unlike prior approaches based on generalizing priority orderings, our policy is not priority-based but rather is designed to induce locally linear mean-field dynamics.

Tue 29 SeptMachine Learning
The gist
This paper looks at a complex decision-making problem where many smaller units, called arms, share limited resources while making choices. Previous work could only guarantee that the difference between a chosen strategy and the best strategy shrinks slowly as the number of units grows. The authors identify specific conditions where this difference shrinks much faster, and they design a new strategy to achieve this better rate. Unlike past methods that rank choices by priority, their approach uses a smoother, linear dynamic to guide decisions.
Open → 2609.38132v1

HelixWorld creates real-time worlds with matching sounds and visuals

HelixWorld: A Real-time Interactive Audio-Visual World Model

Abstract: World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Worlds in video games and simulations usually show pictures and movement but don't add matching realistic sounds that change as you move around. The authors build HelixWorld, a system where pictures and stereo sounds change together in real time based on where the user looks and moves. They collected special data linking real camera positions with sounds to train their system. Their model runs fast on one computer chip and keeps the audio and visuals perfectly synced as you interact, making the experience more immersive than before.
Open → 2609.38123v1

Kv cache quantization reduces memory for long context language models

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.

Tue 29 SeptMachine Learning
The gist
Long conversations with AI need a lot of memory to remember what was said before, which slows things down. The authors came up with a way called WUSH-KV that shrinks this memory by using smart math based on the data’s patterns, so less space is needed without losing much information. They tested it on different language models, and it worked as well or better than other methods even when squeezing the memory to just 2 bits of data. This makes it easier to run AI models on long texts efficiently.
Open → 2609.38121v1

Stochastic world models improve verifying vision-based neural systems

Stochastic World Models for Verifying Vision-Based Neural Feedback Systems

Abstract: Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames more faithfully than GAN surrogates with up to 130 times as many parameters. To verify these surrogates, we develop a procedure that combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark with a GAN surrogate, our procedure resolves the entire state space, 38% of which the state-of-the-art verifier left unresolved. On the RGB version of the benchmark, where no verification results have previously been reported, our procedure resolves over 80% of the state space with a world model surrogate.

Tue 29 SeptArtificial Intelligence
The gist
Verifying systems that use camera images with neural networks is hard because the models of what the cameras see are often too big or inaccurate. This paper shows a new way to model the camera observations using stochastic world models that simulate sensor variations more realistically but remain easier to analyze. The authors also introduce a method combining different verification strategies that can check much more of the system’s possible states than before. Their method works better than previous approaches on a safety test involving emergency braking in driving scenarios.
Open → 2609.38120v1

VideoLoop improves long video understanding by managing memory better

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

Abstract: Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Understanding long videos needs AI systems to think step-by-step, remembering important details while ignoring distractions. The authors found that simply adding more memory can cause the system to get confused by unrelated information. They created VideoLoop, which uses two memory loops to keep track of key points and rewrite its memory to stay focused. This approach helps AI agents answer questions about long videos more accurately than before.
Open → 2609.38119v1