Week beginning 28th September 2026

Every computer science paper posted to arXiv this week, with plain-language summaries and practical uses for each one. Includes commercial applications where relevant.

FurE speeds up realistic 3D animal fur reconstruction by ten times

FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets

Abstract: Realistic and editable animal fur reconstruction from multi-view images is challenging due to fine-scale detail, self-occlusion and obfuscation, and, unlike human hair, the lack of animal-fur datasets. Fur usually covers most of an animal's body, with large inter-species and intra-species variability. We present FurE, an efficient strand-based animal fur reconstruction method that recovers a per-strand, editable groom by optimizing a root-conditioned latent field, decoded into strand geometry via a PCA-based decoder. We reconstruct a defurred animal body using local fur-thickness cues from a surface-constrained Gaussian Frosting representation together with part-based priors. We further show that a PCA-based decoder learned from human-hair strand data can alleviate animal-data scarcity while enabling substantially faster optimization. FurE achieves a 10x speedup in strand training over current SOTA dense per-strand optimization while retaining strand fidelity and generalizing across synthetic and real-world sequences, with quantitative and qualitative validation despite the reduction in training time.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceGraphics
The gist
Creating detailed 3D models of animal fur from pictures is hard because of tiny hairs, overlapping, and no big sets of animal hair pictures to learn from. The authors developed FurE, a way to rebuild each hair strand efficiently by using a kind of model trained on human hair data instead of animal hair data. FurE can quickly create editable, detailed fur that matches the real animal without needing a lot of training time. This method works fast and well, even on both computer-generated and real animal images.
Open → 2609.35770v1

Language models trained to run well at any compute budget

Telescopic Language Models

Abstract: One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.

Mon 28 SeptComputation and LanguageArtificial Intelligence
The gist
Language models usually need different training or compression steps for different amounts of computing power. The authors trained a single model, called a telescopic language model, that works well whether it runs a little or a lot. They do this by training the model so that any smaller version of it still predicts language reliably without needing changes at use time. This approach reduces training costs and provides a smooth range of compute-versus-performance options. Their results show the key is the training method, not just the model’s design.
Open → 2609.35769v1

Video diffusion models improved with projected distribution matching

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Abstract: Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.

Mon 28 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Video diffusion models create videos by gradually refining noisy images, but this process can be slow and sometimes produces poor quality or unnatural effects. The authors identify that errors in a key part of the training process called the critic cause these problems to build up over time. They introduce a new method, PDMD, that filters out these errors to improve stability and video quality without adding complexity or extra steps. Their method works better than previous approaches in tests, making shorter video generation processes produce clearer and more natural results.
Open → 2609.35768v1

Unified models learn to improve image edits by self-feedback loops

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Abstract: Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Creating better images with AI is hard because fixing mistakes often requires trying and checking results repeatedly. The authors show how to train one AI model to not only make images but also to look at its own outputs and decide how to fix them using a special type of learning called reinforcement learning. Their method lets the model learn from multiple correction steps together, making it better at improving images over time without needing extra checking tools. This approach led to better image editing performance on several tests, including ones the model never saw before.
Open → 2609.35767v1

Computational method finds biblical references in karen blixen stories

Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales

Abstract: Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.

Mon 28 SeptComputation and Language
The gist
Finding hidden references to the Bible in literature can be hard, especially when the text doesn't quote directly but hints or paraphrases. The authors worked on spotting these references in Karen Blixen's Seven Gothic Tales by comparing story passages to old Danish Bible translations. They tested different methods and improved one by training it with examples, which helped spot many more subtle references. Their approach serves as a helpful tool to suggest possible connections for scholars to explore further, rather than replacing expert reading.
Open → 2609.35765v1

Reliability-gated fusion improves lower-body 3d pose from consumer imu devices

Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose

Abstract: Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift or drop out. On a new 35-take single-subject benchmark pairing an earbud head inertial measurement unit (IMU) with two smart-insole foot IMUs (SAM-3D-Body pseudo-ground-truth labels), we show the reliability problem is channel-level: a channel ablation isolates foot acceleration as the most informative input (66.6 mm vs. 79.0 mm head-only) and the firmware-fused foot orientation as the liability that destroys the gain. We therefore let the model learn how much to trust each channel of each stream: one temporal gate per stream per channel block, trained with an auxiliary reliability objective on synthetically corrupted pretraining data. The channel-gated model is the most accurate of our learned fusion arms on clean data (69.4 mm vs. 83.7 static, 86.6 ungated) and under every simulated fault (bias in training; drift, dropout eval-only); its gates suppress the natively biased foot-orientation channels on clean real data without test-time supervision and flag dropout bursts at 0.92-0.999 AUROC. Two contrasts: dropping a channel known a priori to fail is flat across foot faults but collapses when an unanticipated stream fails (head dropout: 92.9 vs. 79.3 mm); and a fine-tuned HMD-Poser is more accurate on clean data (64.4 mm) and nominally under drift, with no significant paired difference under bias or dropout, but a larger worst-case degradation from clean (+16.1 vs. +3.5 mm, single seed). Learning to gate reliability instead of sensor count is the lever for deployable sparse inertial capture. Code is available at https://github.com/ZhilinGuo/reliability-gated-imu-fusion.

Mon 28 SeptComputer Vision and Pattern RecognitionHuman-Computer Interaction
The gist
Measuring body movements without cameras can be tricky with common sensors found in everyday gadgets because these sensors sometimes give unreliable data. The authors studied how to better combine signals from a head sensor (like in earbuds) and foot sensors (smart insoles) to capture leg movements more accurately. They found some data from the foot sensors were often misleading and so taught a model to decide which sensor signals to trust at different times. This method made pose estimates more accurate and robust, even when sensors fail or drift over time. Their approach can help make motion capture from simple sensors more dependable and practical.
Open → 2609.35764v1

A new method improves one-step AI image generation accuracy

Unifying Distributional Training for One-Step Visual Generation

Abstract: \emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/

Mon 28 SeptMachine Learning
The gist
Generating realistic images in one step is hard because the system must match many complex details all at once. The authors propose a new approach called MGFlow that better captures and compares complex image features by grouping them into flexible clusters. This helps the model avoid common problems like focusing too narrowly on certain details and produces much better images on a standard big dataset. They also show it can improve text-to-image models to create good pictures faster.
Open → 2609.35763v1

Mobile robots learn complex two-hand tasks from human whole-body videos

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

Abstract: Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.

Mon 28 SeptRobotics
The gist
Teaching robots to use both hands while moving around is very complicated because it needs smooth coordination of walking, whole-body movement, and finger movements at the same time. The authors created DexRoam, a system that learns these skills by watching humans perform tasks while wearing simple VR gear, without needing extra cameras or trackers. They carefully align the human movements to robot actions to keep important details intact. Their tests show that adding human demonstrations significantly improves robot learning, cutting down the number of robot trials needed and boosting success rates.
Open → 2609.35761v1

Llms keep story details straight when writing very long novels

Scaling Long-Form Story Generation via Narrative State Tracking

Abstract: LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories of up to about ten thousand words, leaving their ability to scale to full-length novels underexplored. In this work, we introduce Narrative State Tracking Agent (NstAgent), a training-free agentic framework that allows LLMs to track a structured narrative state including characters, past events and future requirements. We extend an existing benchmark to compare narrative consistency across lengths, and use it together with a writing-quality benchmark to systematically evaluate stories ranging from 10K to 100K words. We show that NstAgent achieves better narrative consistency and writing quality as stories grow longer, and neither of them degrades noticeably as length increases, suggesting that it provides an effective approach to scaling story generation toward full-length novels.

Mon 28 SeptComputation and Language
The gist
Writing very long stories like full novels is hard for language models because they often forget important details. The authors present a method called Narrative State Tracking Agent (NstAgent) that helps these models keep track of characters and events as they write. This method improves the consistency and quality of stories even when they get much longer, up to 100,000 words. It does this without needing extra training. Their tests show that stories stay coherent and well-written as they get longer with this approach.
Open → 2609.35759v1

Tokencast predicts token use during llm tasks to cut waste

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Abstract: When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.

Mon 28 SeptMachine LearningArtificial IntelligenceSoftware Engineering
The gist
When large language model agents solve problems, the number of words they need to use can change a lot each time. This makes it hard to guess how much "talking" (token use) will be needed before they finish. The authors made TokenCast, a tool that watches each part of the task and learns how much words it costs. It then adds these costs up to make good guesses, updating them as the task goes. This helps save on excess word use without extra work from the language model itself.
Open → 2609.35760v1

Adaptive control method predicts and counters dynamic disturbances accurately

Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control

Abstract: We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-layer adaptive disturbance-rejection methods. Specifically, we introduce a statistically principled hard expectation-maximization (hard-EM) procedure, with a Kalman smoother in the hard E-step, to identify dynamical representations of disturbance whose latent evolution is uniformly contractive. The learned representation evolves a latent disturbance-excitation state from measured plant features and control inputs and decodes that state into the time-varying disturbance acting on the nominal plant, thereby extending prior "fixed-decay" last-layer adaptive methods to a learned, predictive DAC-style formulation. Combined with Bayesian filtering of the learned latent state, this representation yields a composite adaptive tracking controller with predictive capability and provable exponential convergence to a bounded neighborhood. We validate our approach experimentally on a slippery ground vehicle carrying a liquid-sloshing tank and a pendulum load, and we further assess its robustness on a system of coupled Duffing oscillators. Across both settings, the method achieves accurate disturbance prediction and improved overall tracking performance relative to fixed-decay representation-learning ablations, LTI disturbance-accommodating baselines, and model-based PD baselines.

Mon 28 SeptMachine LearningRobotics
The gist
Controlling machines that face unpredictable forces, like a vehicle on slippery ground or a pendulum swinging, is hard because these forces can change and affect performance. The authors developed a new method that learns how these disturbances evolve over time by analyzing sensor data and control signals, helping the controller predict and adjust better. This approach improves control accuracy and robustness by combining classic disturbance handling techniques with modern learning-based adaptations.
Open → 2609.35758v1

Neural network solves elliptic PDEs on changing 3D shapes efficiently

Neural Harmonic Measure Operator

Abstract: We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes the density of this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary values on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.

Mon 28 SeptMachine Learning
The gist
Solving certain math problems called elliptic partial differential equations (PDEs) on shapes that can change is hard and usually slow. The authors created a new neural network tool that learns how to predict solutions based on the shape's geometry alone, without needing to retrain for different conditions. This makes it faster and more flexible than past methods. It works well on 3D shapes that vary and can handle different inputs like shapes and applied forces at once.
Open → 2609.35752v1

Looped mixture of experts improves model efficiency by reshaping experts

How to Loop MoE: Flatten the Experts, Untie the Attention

Abstract: Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.

Mon 28 SeptMachine LearningArtificial IntelligenceComputation and Language
The gist
Large AI models sometimes reuse parts of themselves repeatedly to do more with less. This paper shows a way to reshape and reorganize these reused parts, called experts, so the system can choose from a bigger and better pool each time. The authors found that this approach leads to better learning and more balanced use of the parts, without increasing the model size or computation needed. This means AI models can perform better by simply looping through a clever arrangement of their own components.
Open → 2609.35751v1

KV-streams boost training speed for long-context agentic AI models

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Abstract: Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.

Mon 28 SeptMachine LearningArtificial Intelligence
The gist
Large language models that act like agents need to remember long conversations or experiences, but this uses a lot of GPU memory and slows training down. The authors propose KV-streams, a way to keep memory use steady and speed training by streaming data forward instead of restarting it all the time. This makes training 2.6 to 5 times faster without hurting how well the model learns. They also found that this streamed memory helps the model remember important information for longer, improving its decision-making.
Open → 2609.35750v1

Language agents that communicate effectively with fewer words

Towards Communication-Efficient Social Intelligence in Language Agents

Abstract: Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.

Mon 28 SeptComputation and Language
The gist
Socially smart language agents need to cooperate by sharing just the right amount of information, so conversations stay useful without being long or confusing. The authors created a method called Teacher-Assisted Communication Training (TACT) that teaches these agents to pick their words and actions carefully, helping them meet goals while using fewer words. TACT works by checking different ways of saying something, guessing how the partner might respond, and then learning which version works best with the smallest cost in words. The authors tested TACT in two settings and found it improved success rates while cutting down on unnecessary communication.
Open → 2609.35749v1

Looped transformers improve accuracy scaling with adaptive iterations

Improving Test-Time Scaling with Adaptive Looped Transformers

Abstract: Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.

Mon 28 SeptComputation and LanguageMachine Learning
The gist
Longer outputs in transformer models require more computing to stay accurate. The authors studied looped transformers that reuse the same layers multiple times to save parameters but noticed many tokens don’t benefit equally from repeated processing. They created TaH2, which adaptively decides which tokens need extra attention, improving accuracy without extra costs. This method made models more efficient and precise especially as output length and iteration depth increase.
Open → 2609.35748v1

Near optimal limits found for testing balanced parentheses and string equality

Near-Optimal Bounds for Testing Residual-String Equality and Parenthesis Languages

Abstract: Residual-String Equality, denoted $\texttt{ResStringEq}$, is the property consisting of all pairs of strings over $\{0,1,*\}$ that are equal after deleting all `$*$' symbols from them. This property was first introduced by Fischer, Magniez, and Starikovskaya (SODA 2018), who used it to show a lower bound on testing the $\texttt{Dyck}$ languages, where $\texttt{Dyck}_m$ is the language consisting of balanced sequences of parentheses over $m$ parenthesis types. They showed that testing $\texttt{ResStringEq}$ on inputs of length $n$ requires $Ω(n^{1/5})$ queries, and presented a reduction from testing $\texttt{ResStringEq}$ to testing $\texttt{Dyck}_m$ where $m \geq 2$. Furthermore, they showed that $\texttt{Dyck}_m$ can be tested with $O(n^{2/5+δ})$ queries for every constant proximity parameter, where $δ>0$ is an arbitrarily small constant. In this work, we nearly close the remaining gap, by showing that testing $\texttt{ResStringEq}$, and hence $\texttt{Dyck}_m$ where $m\geq 2$, requires $Ω(n^{2/5})$ queries. We also show a stronger lower bound of $Ω(\sqrt{n})$ for testers that make non-adaptive queries. We establish that the $Ω(\sqrt{n})$ bound is nearly tight, by presenting a non-adaptive tester for $\texttt{ResStringEq}$ that uses $O(n^{1/2+δ})$ queries for an arbitrarily small constant $δ>0$. Furthermore, we extend this non-adaptive tester to the $\texttt{Dyck}$ languages, with the same query complexity. Finally, we improve the dependence on the proximity parameter $ε$ in the tester of Fischer, Magniez, and Starikovskaya, reducing it from $O(1/ε)^{\mathrm{poly}(1/δ)}$ to $O(1/ε)^{O(\log(1/δ))}$.

Mon 28 SeptData Structures and Algorithms
The gist
The paper studies how hard it is to check if two special strings match when ignoring certain placeholder symbols, and if a sequence of parentheses is properly balanced. The authors sharpen previous estimates, showing nearly exact boundaries on the number of checks needed to test these properties efficiently. They also provide better methods for testing these conditions without changing their strategy based on earlier answers, improving speed and precision. Additionally, they reduce the complexity of how the testing depends on how close an input must be to being correct.
Open → 2609.35746v1

Linear vision transformers close gap with pretrained softmax weights

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Abstract: Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
The gist
Training new types of vision models called linear ViTs is usually hard and less effective compared to traditional ones that use softmax attention. The authors show that simply copying the attention parts from pretrained softmax models doesn't work well because those parts behave differently. Instead, copying the other parts (called MLP weights) and teaching the attention in a new way helps linear ViTs perform just as well or better. This method works for different model sizes and datasets, making it easier to reuse existing models when switching to linear attention.
Open → 2609.35745v1

Finance agents get automatic expert-guided evaluation rubrics

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Abstract: Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts' key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.

Mon 28 SeptArtificial Intelligence
The gist
Evaluating financial research agents needs careful and expert-based rules to judge their answers accurately. The authors created FinAutoRubric, a system where experts give general guidelines that computer agents use to build detailed, specific rules for checking answers. These rules help write and review financial questions automatically, with human checks for problems. The system was tested on real finance questions and matched expert scores well, even earning preference from analysts in blind tests.
Open → 2609.35744v1

InfiniHand streams accurate hand motion tracking from first-person video

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video

Abstract: World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking how hands move in 3D from video recorded by a camera worn on the head is hard because the camera moves too. The authors present InfiniHand, a way to estimate a hand’s 3D shape and position along with the camera’s motion all at once, without needing separate tools for each. They trained InfiniHand with a very large set of videos and showed it works better and faster than previous methods. This makes it easier to follow hand movements accurately even while the camera wearer is moving around.
Open → 2609.35743v1

Measurement driven modeling reveals unique generative AI network traffic patterns

MINT: Modeling GenAI Impact on Network Traffic

Abstract: Generative AI (GenAI) is becoming a mainstream network workload, yet packet-level simulators lack measure\-ment-driven GenAI traffic models. Currently researchers must approximate GenAI services using traditional sources such as file transfer and video streaming, limiting realistic network evaluation of scheduling and capacity planning. We present MINT, a measurement and modeling framework for GenAI network traffic. Using an isolated net\-work-namespace capture pipeline, we collect client-side traces from three LLM providers across four modalities, cloud and edge servers, and wired and wireless network access points. We find that GenAI modalities exhibit distinct upload/download asymmetry and burst structures that differ from traditional applications. MINT clusters and models these burst regimes and validate empirical burst timing distribution behavior in ns-3 with normalized Wasserstein distances of 2--25\%. Our results also reveal that realistic packet bursts have significantly more variability than constant token generator models. MINT open-sources the first measurement-driven GenAI traffic model for packet-level network simulation.

Mon 28 SeptNetworking and Internet Architecture
The gist
Generative AI, like chatbots and image generators, uses networks differently than older internet services. The paper presents MINT, a way to measure and simulate these new traffic patterns based on real data from multiple AI providers. It shows that AI traffic uploads and downloads are uneven and come in bursts that standard models don’t capture well. This new model helps networks prepare for and manage the growing AI-generated data flow more realistically.
Open → 2609.35742v1

Self-explaining language models improve task solving without reinforcement learning

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Abstract: People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

Mon 28 SeptArtificial IntelligenceComputation and Language
The gist
People learn not only by doing but also by thinking about and explaining their experiences, which helps them improve. The paper shows that a language model can train itself to do better on tasks just by explaining what it did, without needing rewards or teachers. This method, called ROFT, made the model solve more problems and even learn from tasks where it initially failed every attempt. The explanations help the model figure out which actions were good or bad, leading to better future behavior. This suggests teaching AI to explain could help it learn more effectively.
Open → 2609.35741v1

Rubric calibration improves accuracy of large language model document rankings

Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory

Abstract: Rerankers decide which documents users and LLMs see, yet their standard metric, nDCG, relies on human relevance labels that are costly, sparse, noisy, and discretely graded. As rerankers approach each other in quality, nDCG on these labels therefore increasingly fails to separate them. LLM judges could supply dense labels. Relative judgments within one query tell even close candidates apart, yet their scores share no scale across queries. Absolute grades share one scale but are too coarse to distinguish documents of similar relevance. We propose Rubric-Calibrated Preferences (RCP), which combine both kinds of judgment. A listwise Bradley-Terry tournament orders each query's documents, and a rubric of yes/no criteria of increasing stringency provides an absolute standard. Item Response Theory (IRT), which scores test-takers based on their answers to common questions, then uses the shared criteria to put all queries' tournament scores on one scale. RCP's retrieval metric, RCP-nDCG, replaces nDCG's discrete labels with the resulting calibrated relevance probabilities. Against blind grades from 46 external annotators, calibration raises the correlation between a query's mean score and its mean human grade from 0.538 to 0.795. The probabilities rank a useful document above a non-useful one with probability 0.910 (AUC, chance 0.5), versus 0.651 for the benchmark labels. When the annotators' grades prefer one of two rerankers and exactly one metric agrees, that metric is RCP-nDCG in 72.4% of 185 comparisons (chance about 53%). On TREC-DL, RCP-nDCG sides with NIST assessors' grades on every reranker pair that these grades separate significantly. RCP-nDCG also resolves many of nDCG's ties and separates 1.9 times as many reranker pairs on NanoBEIR. Rubric calibration thus turns relative LLM judgments into dense relevance labels that are comparable across queries and agree with human judgment.

Mon 28 SeptInformation Retrieval
The gist
Deciding which documents show up first in search results is tricky because human labels that judge relevance are limited and inconsistent. The authors created a method called Rubric-Calibrated Preferences that combines detailed comparisons within searches and clear yes/no standards across searches to better rate document relevance. This method uses a testing theory to put all these ratings on the same scale, improving agreement with human judgments and distinguishing closely ranked results more effectively. Their new metric, RCP-nDCG, better separates good documents from less useful ones compared to traditional methods.
Open → 2609.35739v1

Language models learn to improve their task strategies at run time

Harness Learning Enables Generalizable Test-Time Adaptation

Abstract: A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.

Mon 28 SeptComputation and LanguageMachine Learning
The gist
Modern AI language models work with instructions called harnesses that tell them how to organize their actions and tools for different tasks. This paper presents a way for the model to learn how to improve these instructions by trying them out and getting feedback, without changing its internal model itself. This lets the model adapt to new tasks it hasn’t seen before by revising how it solves problems during use. The authors show that this approach helps the model get better at reasoning and answering questions that require multiple steps.
Open → 2609.35738v1

GeoVerse improves world-consistent new views from sparse images

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Abstract: Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Creating new views of a scene from a few images is hard because the computer must guess what the unseen parts look like and keep things consistent as the viewpoint changes. The authors designed GeoVerse, which works inside a 3D model’s brain-like space and borrows ideas from video generators to fill in missing parts better. It also remembers everything seen so far to keep the whole scene aligned when making new views. Their experiments show GeoVerse creates clearer and more reliable pictures than earlier methods.
Open → 2609.35734v1

Failure-transparent agents reduce false success claims in AI tools

Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

Abstract: Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.

Mon 28 SeptArtificial Intelligence
The gist
Sometimes AI agents use tools that fail, and then the agents incorrectly say they succeeded without proper proof. The authors created a special test called Failure-Transparent Agents (FTA) to focus on checking whether agents honestly report failure after tools fail. They tried different models and methods, showing that adding clear instructions and requiring structured evidence drastically reduces false success claims and made responses more helpful. This work helps improve trustworthiness in AI that uses other tools.
Open → 2609.35732v1

Interactive humanoid videos created with real time multimodal control

FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

Abstract: We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video backbone to accept rolling action prompts, streaming audio, and dynamically updated image, audio, and video references. Video-driven rotary positional embeddings align reference chunks with the generation timeline, while reference-plus-image conditioning and partially noised historical motion frames preserve appearance and avoid accumulated drift. Second, a Proactive Interaction Agent separates pre-online planning from online scheduling and response: it prepares a persona, a long-horizon agenda, and reusable multimodal skills in advance, then autonomously schedules behaviors, responds to audience input, and handles interruptions during a live session. FlowAct-R2 supports real-time 720p generation and hour-scale streaming across entertainment streaming, live shopping, video chatting, and live vlogging.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Creating realistic animated videos of people that interact in real time is a challenge because it requires understanding and reacting to many inputs like speech, images, and actions. The authors introduce FlowAct-R2, a system that uses a special technique to take in ongoing audio and images while planning ahead how the animated person should behave. This helps the video avatar maintain a consistent look and respond smoothly to what viewers say or do live. The system can create high-quality videos that run continuously for a long time and fit scenarios like live streaming, shopping, or chatting.
Open → 2609.35728v1

Camera angle strongly affects AI quality in rehab exercise monitoring

Impact of Patient Orientation in Single- and Multi-View Camera Environments for AI-based Rehabilitation Monitoring

Abstract: Automated quality assessment of rehabilitation exercises relies heavily on accurate human pose estimation from video data. Although numerous RGB-based pose estimation methods have been proposed, the impact of camera placement on detecting clinically relevant movement errors remains insufficiently explored. To address this gap, we introduce REHAB26-ViewAngles, a dataset comprising correct and incorrect rehabilitation exercise executions captured from a wide range of camera angles. Furthermore, we propose a novel separability metric to quantify an algorithm's ability to distinguish between valid and faulty exercise repetitions. Using these tools, we analyze how various RGB-based pose-estimation strategies are suitable for exercise quality assessment under varying camera placements. In particular, we analyze single-camera 2D and 3D pose estimation and four multi-camera strategies: a combination of two orthogonal 2D views, 3D triangulation, weighted 3D fusion, and an AI-based pose-estimation transformer model specifically trained from two synchronized cameras. Our findings reveal that an optimally placed 2D camera can improve the separability by 16.9\,\% over the commonly used $0^\circ$ frontal view and frequently outperforms single-camera 3D estimation, while combining two views can further improve accuracy by up to 13.1\,\%. These results offer practical guidance for deploying rehabilitation monitoring in both home and clinical settings.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking how well people do rehab exercises with AI depends a lot on where the camera is placed. The authors made a dataset with many angles showing both good and bad exercise attempts to study this. They created a way to measure how well AI can tell right from wrong movements depending on camera views. They found that picking the best single camera angle can improve error detection more than some 3D methods and that using two cameras gives even better results. This helps decide how to set up cameras for rehab monitoring at home or in clinics.
Open → 2609.35726v1

Geometric aware superquadric fitting improves 3d shape decomposition

Superquadric Primitive Decomposition of 3D point clouds via Geometric-Aware Inlier Refinement

Abstract: The decomposition of 3D point clouds into interpretable geometric primitives remains a longstanding challenge in Computer Vision and Computer Graphics. Among the available representations, superquadrics offer a compact and expressive model capable of capturing a wide range of shapes. However, their estimation is inherently challenging, as it requires solving a non-linear optimization problem and is particularly sensitive to noise, outliers, and overlapping structures. While robust estimation methods such as RANSAC and its variants achieve strong performance, they rely primarily on spatial proximity and residual-based criteria, often leading to incorrect inlier assignments across adjacent or complex arrangements of primitives. In this work, we introduce a geometric-aware framework for primitive decomposition that explicitly incorporates local surface properties into the fitting process. Specifically, we propose an inlier refinement step formulated as an energy minimization problem and solved via graph-cut optimization. Our formulation integrates geometric priors, such as normal consistency, enabling more reliable inlier selection beyond purely residual-based criteria. The approach naturally applies to both single-model estimation and multi-model decomposition. By leveraging geometric information beyond point-wise residuals, our method reduces erroneous inlier propagation and stabilizes parameter estimation. Experiments on synthetic and real datasets show consistent improvements in geometric accuracy, robustness to noise and outliers, and convergence efficiency compared to state-of-the-art RANSAC-based methods.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Breaking down 3D point clouds into simple shapes helps computers understand objects better, but it's tricky due to noise and overlapping parts. The authors introduce a new method that not only looks at how close points are to a shape but also checks surface details, which helps find better matches. They use a math trick called graph-cut optimization to refine which points belong to each shape, making the process more accurate and stable. This approach works well for finding one or many shapes and is better than older methods on both fake and real data.
Open → 2609.35725v1

AI agent swarms make steady progress in scientific research tasks

AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions

Abstract: Artificial intelligence (AI) agents, language models connected to tools and run in a loop, can now carry out long, multi-step tasks with little supervision. We gave swarms of off-the-shelf coding agents a short statement of scope, from a narrow topic to a whole field, access to the literature and to computing tools, and one standing instruction: make real, correct, useful progress, and do not stop. We supplied no scientific ideas. Within weeks, the agents produced a large body of research notes, paper-length drafts, and formal proofs in five areas of optimization theory and physical science, and proposed untested laboratory experiments in a sixth. We do not claim that all of it is correct or new, but it is not noise: in what we have checked so far, we found no major scientific error, and several results are proved in a proof assistant. The agents produced results faster than we could review them; we estimate that a full review would take us months. Together with two widely discussed 2026 results in mathematics obtained with swarms, our runs suggest that agents can already do a large part of routine theoretical research, at least in areas that we experimented with. This raises questions we cannot yet answer: how to trust results when review, not production, is the scarce resource; what credit and publication counts mean when the human input is a prompt, and why institutions would pay researchers rather than buy computing time; and how people can learn a field, add to what agents do, and stay in control of research they cannot keep up with. Research institutions are not ready: models improve faster than institutions change, so they should decide now how to respond as capabilities increase. We offer tentative positions, release the agents' unedited output as of 25 September 2026, and invite readers to repeat the experiment in their own fields.

Mon 28 SeptComputers and Society
The gist
AI programs working together can now carry out complex research tasks without much human help. The authors gave groups of these AI agents access to scientific papers and tools, with only a simple goal: make real progress. In a few weeks, the agents produced lots of research notes, drafts, and proofs in math and science fields. While not all of it was checked, what was reviewed had no major errors, showing these AI swarms can do routine theoretical research faster than humans. This raises important questions about trusting, reviewing, and crediting work done by AI.
Open → 2609.35719v1

GPT 6 Astra narrows gaps in semantic vision but struggles with precise tasks

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Abstract: Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Understanding images is becoming easier for general AI systems like GPT-6 Astra, which can do many challenging vision tasks previously handled only by specialized models. The authors found that Astra excels at understanding the meaning of scenes and reasoning about objects but still struggles with tasks that require exact measurements, detailed reconstructions, or consistency over time in videos. Some additional tools help improve performance on specific tasks, but gaps remain in high-fidelity visual perception. This work shows where general AI is getting good at vision and where specialized efforts are still needed.
Open → 2609.35718v1

RoboCompiler simplifies consistent robot modeling and control

RoboCompiler: Graph-Native Compilation of Closed-Chain Robots for Consistent Modeling, Control, and Simulation

Abstract: Robots with kinematic loops, coupled actuators, and changing contacts require consistent models of configuration, motion, force, and dynamics. Yet these interfaces are often reconstructed separately for control and simulation, making closure and actuation consistency difficult to maintain. This paper presents RoboCompiler, a graph-native framework that compiles a canonical mechanism graph into a shared mechanical interface. From bodies, joints, frames, inertias, and actuator ports, it constructs closure paths and analytic residual Jacobians, then assembles feasible configurations through rank-checked continuation and correction. A tangent lift maps independent velocities to full robot and task motion, while paired actuator-port maps preserve virtual work. A constraint-curvature correction extends the reduction to accelerations and projected rigid-body dynamics, including floating-base and support modes. Cycle-local evaluation, generated Jacobians, and dependency-aware reuse enable localized updates when closure inputs change. We evaluate physical loops and task-induced constraints on a Komatsu excavator, Unitree Go2, Franka Panda, Kangaroo, and a six-UPS Stewart platform. High-precision constrained-dynamics and independent Pinocchio checks confirm mechanical consistency; MuJoCo and Isaac Sim/PhysX executions demonstrate task performance and model reuse under native contact. For Kangaroo, compilation reduces residual-and-Jacobian evaluation time by 96.7% and closed-loop rollout wall time by 66.8%, with dynamics and control held fixed.

Mon 28 SeptRobotics
The gist
Robots with complex moving parts, like loops and linked joints, need precise models that work the same for controlling and simulating them. The authors created RoboCompiler, which turns a robot's mechanical setup into a shared model that keeps everything consistent across different uses. This approach speeds up calculations and keeps the robot's motions and forces accurate in simulations and real controls. The system was tested on real and simulated robots, showing it works well and reduces computational time significantly.
Open → 2609.35717v1

X-Reset trains robot hands to grasp diverse objects using human resets

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Abstract: Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.

Mon 28 SeptMachine LearningArtificial IntelligenceRobotics
The gist
Training robot hands to pick up and use many different objects is hard because robots struggle to learn these skills from scratch. The authors propose X-Reset, a method that uses snapshots from human hand motions interacting with objects to help robots start their learning in helpful positions. Instead of copying human movements exactly, X-Reset converts these human poses into robot poses and carefully chooses stable states to reset to during training. This approach allows robots to learn general skills for many objects across different types of robot hands, even working with imperfect human data and transferring skills from simulation to real robots.
Open → 2609.35715v1

Competitive betting improves test power but limits expected gains

Competitive optimality in testing by betting via Bell-Cover randomization

Abstract: Bell and Cover showed that an investor who multiplies the initial unit of capital by an independent uniform random variable on $(0,2)$, and then uses the log-optimal portfolio, wins a head-to-head wealth comparison with probability at least one half against every independently randomized competitor. We explain very simply how this result transfers to testing by betting: for any composite null $\mathcal P$ and simple alternative $Q$, denoting $E^*$ as the corresponding numeraire e-variable, we show that $UE^*$ exceeds any other e-variable $E$ with probability at least half. Interestingly, we show that this competitive optimality result is actually equivalent to the numeraire inequality $\mathbb E_Q[E/E^*]\leq1$, and in general randomization only helps the numeraire and fails to improve the competitive advantage of an arbitrary e-variable. Under optional stopping with or without knowledge of $U$, we emphasize a key distinction between e-process validity and competitive optimality. We also show that competitive optimality comes at the price of expected log wealth and power: thresholding $UE^*$ at $1/α$ has sharp size at most $α/2$, but the factor of two actually disappears under optional stopping. Even after correcting for this factor of two, the test is dominated in conditional rejection probability by randomizing the testing threshold (randomized Markov's inequality). Thus, Bell-Cover randomization is optimal for a specific competitive objective, at the cost of others.

Mon 28 SeptComputer Science and Game TheoryInformation Theory
The gist
The paper studies a way to improve statistical tests by randomizing betting amounts, inspired by a finance method that beats opponents half the time. The authors show this approach can outperform other methods in certain competitive settings but at a cost to the average success of the test. They also explain when and how randomizing bets helps or doesn't, especially when stopping tests early. This work clarifies trade-offs between maximizing test power and maintaining strong overall performance.
Open → 2609.35714v1

Improved bounds on simple auctions for selling multiple items

On the Power of Determinism in Multi-Item Auctions

Abstract: We study the classical multi-item monopoly setting with a single additive buyer and $m$ heterogeneous items whose values are independent but not necessarily identically distributed. Optimal truthful auctions may be randomized and complicated. We analyze the approximation ratios of three simple deterministic auctions: selling all items separately, selling them as a single grand bundle, and choosing the better of the two. Our technical cornerstone is a nonlinear mathematical programming formulation of the worst-case approximation ratio of selling separately, in discrete auctions where values lie in the grid $\{0,1/K,2/K,\dots ,1\}$. For two iid items, we construct novel tight Lagrangian dual certificates that determine this ratio exactly for any discretization parameter $K$. Taking $K\to\infty$, we obtain the tight bound $1+W(1/e)\approx 1.278$ in the continuous-valued setting, where $W$ denotes the Lambert-W function, closing the $[1.278,1.368]$ gap from the work of Hart and Nisan [EC'12, JET 2017]. For $m\geq2$ independent items, a different dual construction gives an upper bound on the approximation ratio of selling separately in terms of basic statistics of the item values. Combining this bound with new inequalities relating optimal revenue (REV), separate-selling revenue (SREV), and grand-bundle revenue (BREV), we derive improved guarantees for all three auctions. Most notably, we prove \[REV\leq 3.5 \max\{SREV,BREV\},\] improving upon the $5.2$ factor of Ma and Simchi-Levi [AISTATS'21] and the $6$ factor of Babaioff, Immorlica, Lucier and Weinberg [FOCS'14, JACM 2020]. For iid items, we also prove $REV\leq 4.4534 BREV$.

Mon 28 SeptComputer Science and Game TheoryDiscrete Mathematics
The gist
Figuring out how to sell multiple different items to one buyer for the most money can be very complicated. The authors studied three simple, straightforward ways of selling these items: selling each separately, selling all as one bundle, or picking the better of these two choices. They used math to find the worst-case performance of selling separately and improved previous estimates on how close these simple methods come to the best possible income. Their work gives better guarantees that these simple auctions earn at least a certain fraction of the ideal revenue.
Open → 2609.35711v1

Video prediction and image generation improved by physics based model

Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective

Abstract: We introduce LHFM, a geometric framework for learning image dynamics. Drawing on structures central to classical mechanics, symplectic geometry, and geometric quantization, LHFM represents each image as an exact Lagrangian graph and models its evolution through image-dependent Hamiltonian flows, which yield a transport--source parameterization of image velocities. Our primary application is deterministic video prediction: LHFM-V is a recurrent model that advances frames by integrating predicted transport and source fields, and achieves the lowest reported FLOP count among the compared recurrent models with similar prediction accuracy. The image variant, LHFM-I, shows that the same construction is compatible with flow matching: in a matched experiment, it attains a lower FID than the flow-matching baseline.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Predicting future video frames or creating new images can be challenging because it requires understanding how images change over time. The authors use ideas from classical physics to represent images and their motion as mathematical flows, which helps model these dynamics more accurately and efficiently. Their method, called LHFM, uses these flows to predict videos and generate images with better accuracy and lower computational cost compared to some existing methods.
Open → 2609.35710v1