Week beginning 21st September 2026

Every computer science paper posted to arXiv this week, with plain-language summaries and practical uses for each one. Includes commercial applications where relevant.

FuseReg improves image generation by fusing encoder layers flexibly

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
When making computers create images, it's tricky to decide which parts of the visual understanding to use, because some parts help keep fine detail while others improve overall quality. The authors propose FuseReg, a method that trains the system to use different combinations of these parts randomly, so it learns to handle all of them well. This makes the image generator better at both recreating detailed images and producing high-quality new images without needing to change the original visual encoder.
Open → 2609.31620v1

Confidence training improves reasoning efficiency in AI models

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.

Fri 25 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
Long step-by-step thinking by AI models can be very slow and costly. The authors discovered that teaching models to estimate how confident they are in their answers during reasoning can make them think faster without changing how they decide to stop. This method uses only a small set of problems and does not directly encourage shorter answers, yet it reduces the length of reasoning by up to 25% while keeping accuracy the same. The improvement comes from the models learning to be better aware of their own reasoning, not by forcing shorter thought tracks.
Open → 2609.31619v1

Differentially private PCA algorithm works without eigenvalue gaps

Gap-free Differentially Private PCA for Gaussian Data

Abstract: We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.

Fri 25 SeptData Structures and AlgorithmsMachine Learning
The gist
When trying to find the main directions of variation in data (called principal components) while keeping individual data private, existing methods often need clear differences between these directions. This paper presents a way to do this with Gaussian data even when these differences are small or absent. The authors provide a new algorithm that maintains privacy without relying on gaps in the data. This could help analyze sensitive data more accurately and privately.
Open → 2609.31614v1

Reverse diffusions contract divergences and guarantee local stationarity

First-Order Stationarity of Reverse Diffusions

Abstract: Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.

Fri 25 SeptMachine Learning
The gist
This paper studies how certain random processes called reverse diffusions behave when used to generate data. The authors show these reverse processes shrink a measure of difference between probability distributions at a steady exponential rate when the added noise has a particular mathematical property called strong convexity. They also analyze how discretizing these continuous processes still leads to stable results that locally match the correct data patterns. This helps understand how these methods produce realistic samples and relates closely to ideas in optimization.
Open → 2609.31612v1

Post-processing improves generative AI outputs to match target attributes

Statistical attribute alignment for black-box generative AI via output post-processing

Abstract: Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.

Fri 25 SeptArtificial IntelligenceMachine Learning
The gist
It can be hard to make AI-generated content, like pictures or profiles, have specific characteristics that users want, such as fairness or diversity. The authors developed a way to adjust the outputs after they are created, without needing to change how the AI works internally. Their method uses many tries from the AI to pick outputs that better match a desired distribution of features. They tested their method on image and persona generation tasks and showed it complements other ways of guiding AI output.
Open → 2609.31607v1

Robot policies learned from sparse success signals improve task success

Learning Robot Policies from Sparse Success Signals via STL-Guided Stein Variational Policy Gradient

Abstract: Learning robot policies for tasks with sparse success signals is challenging when completion depends on coordinated actions, precise contact outcomes, or satisfying several conditions together. Intricate physical interactions with the world further complicate these requirements. Prior work using conventional reward shaping mechanisms provides dense feedback but local progress might not translate into eventual task completion. We present Signal Temporal Logic-guided Stein Variational Policy Gradient (STL-SVPG), a population-based method that uses smooth STL robustness as a trajectory-level training objective. Differentiating this objective through the dynamics assigns credit to policy actions according to their effect on the complete task specification, rather than local progress alone. We evaluate the approach on six quadcopter and manipulator tasks that involves event-triggered responses, strictly ordered behavior, responses within specified deadlines, and physical interaction with the world. STL-SVPG achieves the highest mean success rate among the compared methods on five of six benchmarks. Simulation-trained policies trained in simulation transfer temporal and contact task behavior to the real world.

Fri 25 SeptRobotics
The gist
Teaching robots to perform tasks can be hard when they get little feedback on whether they succeed or fail. The authors created a new method that helps robots learn better by using a mathematical way to understand how well a whole task was done, rather than just small steps. This method was tested on drones and robot arms doing complex jobs that need careful timing and interaction with the environment. The approach helped the robots succeed more often and even worked when moving from simulation to the real world.
Open → 2609.31606v1

Llms reveal and adjust internal user beliefs to guide responses

User Model Extraction via Belief Self-Distillation

Abstract: Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.

Fri 25 SeptMachine LearningComputation and Language
The gist
Large language models (LLMs) guess information about the person they are talking to and change how they reply based on these guesses. The researchers created a new method called Belief Self-Distillation (BSD) that lets us see and change what the model believes about its user from regular conversations, without needing extra labels. This method helps test how changing these beliefs affects what the model says, for example, why it refuses some requests. They also found that different independently trained models organize their user beliefs in very similar ways. This work helps make the internal thinking of LLMs more understandable and controllable.
Open → 2609.31603v1

LoRA adapters combine smoothly by letting new skills only read old skills

New LoRA Skills Should Read but Never Write

Abstract: Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.

Fri 25 SeptMachine Learning
The gist
Large language models can learn new tasks cheaply by adding LoRA adapters, but merging multiple adapters is tricky because they interfere with each other. The authors found that the way adapters are combined matters a lot: new adapters should only read from existing ones, not change them. Their method, called READ, rewrites adapters into a form that guarantees new skills read but never overwrite old ones. This approach lets models add skills one at a time without losing past knowledge and improves performance on popular benchmarks.
Open → 2609.31600v1

GraphWrit3R generates 3D scene graphs directly from point clouds

GraphWrit3R: End-to-End 3D Scene Graph Writing

Abstract: 3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Understanding 3D environments means recognizing objects and their relationships. Previous methods required complicated steps and extra information that aren’t available in real situations. The authors created GraphWrit3R, a simpler method that takes raw 3D data and directly produces a detailed map of objects and how they relate. It works with different data types and uses a large language model to describe scenes in natural language. Their method matches or exceeds the best existing techniques without needing extra annotations during use.
Open → 2609.31595v1

Automated source code method generates accurate iot network profiles

From Source Code to Network Profile: Automated and Traceable MUD Profile Generation for IoT Devices

Abstract: The Manufacturer Usage Description (MUD) standard allows IoT manufacturers to define expected network behaviors in a MUD file. This file can be translated into enforceable access-control policies, restricting compromised devices to operate solely through manufacturer-defined communication patterns. However, practical adoption of MUD depends on profiles that are accurate, complete, and maintainable. Existing approaches use traffic-based automation but require device deployment and prolonged monitoring, capturing only behavior exercised during observation. Rare, failure-triggered, or configuration-dependent communications may remain absent, producing incomplete policies that disrupt legitimate operation and offer limited insight into the software components responsible for each rule. We present AutoMUD, a source-code-driven tool that generates traceable MUD profiles for IoT devices from their firmware and software source code. AutoMUD combines static and syntactic extraction, retrieval-grounded language-model reasoning, and deterministic validation and compilation to recover the communication behavior characterizing an IoT device and translate eligible endpoints into policy rules. By analyzing code-level evidence, AutoMUD exposes rarely exercised and conditional communication paths, links every generated rule to its source-level provenance, and preserves excluded findings with explicit reasons for review. Our evaluation on a Linux-based repository demonstrates that AutoMUD recovers complete communication behavior, consolidates validated behavior into semantic endpoint groups, and generates structurally valid MUD profiles. Through a controlled semantic fault-injection campaign, we demonstrate that AutoMUD enables analysts to detect, localize, explain, and correct propagated errors, recovering policies semantically identical to their clean counterparts.

Fri 25 SeptCryptography and SecuritySoftware Engineering
The gist
Devices connected to the internet can behave in unexpected ways that make them vulnerable to attacks or malfunctions. The authors created a tool called AutoMUD that reads the actual software code inside these devices to figure out how they communicate on a network. This helps build detailed and trustworthy rules to control device communication and keep them safe. By using code instead of just observing the device in action, AutoMUD finds rare or hidden communication behaviors and links each communication rule back to the exact part of the code that causes it. This makes it easier to understand, check, and fix the network behavior rules.
Open → 2609.31594v1

AgentWorld benchmarks long multi-agent collaboration in games

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Abstract: Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.

Fri 25 SeptMultiagent Systems
The gist
Many current tests for AI agents focus on short or competitive activities and don’t really measure how well multiple agents work together over time. The authors created AgentWorld, which tests teams of AI agents collaborating in a complex online game for many steps and with different roles. They also made a new way to measure how much each agent’s actions contribute to the team's success. When tested on several advanced AI models, none did better than about half the tasks, often failing because agents didn’t communicate well or keep plans. The benchmark and tools are open-source for others to use.
Open → 2609.31590v1

Direct feedback alignment reveals common error collapse slows learning

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Abstract: Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.

Fri 25 SeptMachine LearningNeural and Evolutionary Computing
The gist
Training certain neural networks can get stuck because they share a common error pattern that makes hidden units saturate and stop learning effectively. The authors found that this happens when the network’s hidden units respond too similarly, a problem they call 'common-mode collapse.' They show that adjusting the way the output error is fed back and how learning rates are set can reduce this problem and speed up training. Their study tested different settings, networks, and datasets to understand when and why this collapse happens.
Open → 2609.31589v1

Voice dialogue system plays pre-recorded lines faster with low latency

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

Abstract: Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and playing pre-recorded lines. Using probing, we identify the layer and frame at which the upcoming response becomes recoverable, and use this hidden state as the retrieval query. RePlay retains only the layers up to that point and replaces text and speech generation with lightweight turn-taking and retrieval heads. In simulated multi-turn interviews, RePlay reaches a median latency of 383 ms, 3 to 7 times lower than ASR-LLM cascades of comparable dialogue quality, at the cost of lower exact-line accuracy. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade with a small LLM (p = 0.008), and showed a non-significant preference (46% vs. 21%) over a slower cascade with a stronger LLM.

Fri 25 SeptSound
The gist
Many voice assistants struggle to replay exact, pre-recorded responses quickly during conversations. The authors designed RePlay, a system that retrieves and plays back prerecorded lines with faster response times in multi-turn dialogues. They found a way to predict when a reply is ready to be retrieved early in the model’s processing, so it can quickly fetch the correct audio snippet. In tests, RePlay responded up to seven times faster than traditional systems but traded some accuracy in delivering the exact prerecorded line. Users generally preferred its faster, smoother interaction.
Open → 2609.31588v1

Better documentation does not improve coding agent fixes in repositories

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.

Fri 25 SeptSoftware EngineeringArtificial IntelligenceComputation and Language
The gist
People often think that good natural-language explanations of code help automated coding assistants fix problems better. The authors built a way to test how well code descriptions work by seeing if code rewritten from them still passes tests. They found shorter or longer descriptions don’t matter as much as how complete the explanations are. However, when they tried using improved descriptions to help coding agents fix real issues in software projects, the descriptions did not help. This suggests that having the code and the problem alone is usually enough for these agents.
Open → 2609.31587v1

Trust guided transformer improves long decision making in AI models

Trust Guided Decision Transformer

Abstract: Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.

Fri 25 SeptMachine Learning
The gist
When AI models try to make many decisions in a row, their performance often drops because they rely on past information that becomes less reliable over time. The authors found that the model’s own errors in predicting the next state reveal when this happens. They made a method called Trust Guided Decision Transformer that ignores unreliable past information and uses trustworthy parts to decide the best action. This method improves performance on navigation and movement tasks compared to older approaches.
Open → 2609.31586v1

Humanoid robot improves walking and terrain skills using depth images

Generate, Track, Improve: Perceptive Multi-Skill Humanoid Locomotion with RL-Fine-Tuned Motion Generators

Abstract: General purpose humanoids require locomotion controllers that are multi-skill, perceptive, dynamic, and robust enough to go anywhere humans can. In this work, we present a two layer locomotion architecture: (1) a perceptive flow matching motion generator plans whole body trajectories from raw depth images while a (2) perceptive tracking policy trained with control-guided RL follows these motions. Both policies are trained on a library of terrain consistent motion clips created with dynamically optimized human data which yields both accurate velocity tracking and terrain consistent references. Our central contribution is a simple yet effective off-policy RL fine tuning loop that improves the motion generator. A structured search method is used with the generator to gather data for advantage weighted regression. This off-policy loop is much more sample efficient than on-policy residual fine tuning and improves terrain consistency on unseen geometries and skill compositions. We find that successful terrain traversals increased by up to 25 percentage points and skill selection improved by up to 80 percentage points. By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy. With two cameras, the policy can see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain. A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments. Project page: https://zolkin1.github.io/generate-track-improve/

Fri 25 SeptRobotics
The gist
Making robots walk and move like humans on different surfaces is very hard because they need to see and adjust to the ground quickly. The authors designed a two-part system where one part plans the robot’s whole body movements from camera depth images, and the second part follows those plans exactly. They improved the planning part by letting it learn from its own experiences using a smart data search technique, which helped the robot get better at walking on new surfaces and choosing the right skills. This allows a robot dog-like humanoid to walk, run, jump on boxes, and climb stairs outdoors using just camera input without needing extra maps or sensors.
Open → 2609.31577v1

System prompts mostly manage tools not model ethics

Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts

Abstract: Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. Operational content rather than ethical statements dominates the corpus; a deliberately simple block-level classifier assigns roughly 58\% of classified words to tool/protocol and roughly 5\% to safety policy, while the strictest rule-lines guard tool use and file safety over harmful content by an 11:1 margin. Literal text transfer concentrates in a small set of cross-vendor pairs. Prompts also carry measurable maintenance debt, with version chains turning over thousands of words per release. The evidence supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns. Because most documents are adversarial in origin and the detectors are deliberately simple, all magnitudes are directional; we audit the main classifier's error modes.

Fri 25 SeptCryptography and Security
The gist
People often think leaked instructions for AI language models reveal the models' hidden values or ethics. This study shows that these instructions mostly focus on how to operate tools and keep files safe, rather than on moral guidelines. The prompts also change a lot over time, similar to software configuration, and are reused between vendors sometimes. The authors suggest treating these instructions as technical setups instead of ethical statements.
Open → 2609.31575v1

Improved vertex coloring method reduces imbalance in triangle networks

Polychromatic 2-colorings with Bounded Discrepancy for Triangulations

Abstract: A polychromatic $2$-coloring of a triangulation is a $2$-coloring of the vertices such that no face is monochromatic. The discrepancy of a coloring is the maximum difference between the sizes of the color classes. Asayama and Matsumoto (Graphs and Combinatorics, 2022) proved that every triangulation admits a polychromatic $2$-coloring with discrepancy at most $\tfrac{5n-16}{9}$, and that there exists a class of triangulations for which every polychromatic $2$-coloring has discrepancy at least $\tfrac{n}{3} - 2$, where $n$ is the number of vertices. We improve the upper bound, showing that every triangulation admits a polychromatic $2$-coloring with discrepancy at most $\tfrac{3n-16}{7}$ and such a $2$-coloring can be computed in quadratic time. We also show a discrepancy of at most $n-\tfrac{4M}{3}$ for triangulations with a matching of size $M$. This implies, for example, that Delaunay triangulations admit a discrepancy of at most $\tfrac{n}{3}$. We provide a linear-time algorithm to compute a $2$-coloring whose discrepancy is at most $\tfrac{5n-24}{7}$. One of our results shows that any proper four coloring with the largest color class of size $\frac{n}{2}$ would imply a $2$-coloring with discrepancy at most $\frac{n}{3}$. The existence of such a proper coloring has been recently confirmed by Kawarabayashi, Yoneda, and Yoneda (arXiv 2026). Therefore the two results together confirm the discrepancy of at most $\frac{n}{3}$ for triangulations.

Fri 25 SeptComputational Geometry
The gist
The paper studies a way to color the points (vertices) of a triangle-filled network so that no triangle has all points the same color, using just two colors. The goal is to keep the number of points in each color group balanced, minimizing how uneven the groups are. The authors improved previous results by showing how to achieve a better balance between the two colors and provided efficient algorithms to find these colorings. They also connected their work to special types of triangle networks called Delaunay triangulations and linked their findings to recent results on four-colorings of these graphs.
Open → 2609.31574v1

Inr models improve brain MRI segmentation with fewer parameters

How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI

Abstract: Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight alternative for semantic segmentation, achieving competitive performance with substantially fewer parameters than conventional architectures. However, the mechanisms, scaling behavior, and domain generalization abilities of INR-based segmentation remain insufficiently understood. In this work, we study these questions in the context of cross-domain brain MRI segmentation. We analyze INR-based segmentation across low-parameter regimes, comparing it with conventional pipelines in both in-domain and out-of-domain settings. Surprisingly, we find that INR-based models do not simply improve with increasing parameter budget. Their advantage is most pronounced under low-parameter and limited-augmentation settings, while U-Net-based models benefit more from larger capacity and standard augmentation. We also investigate how INRs encode semantic information in their hidden features and show that complementary segmentation-relevant structure is distributed across multiple INR layers. Building on this insight, we introduce HierINRSeg, a hierarchical INR-based architecture that aggregates multi-layer representations for improved robustness and generalization. Extensive experiments show that HierINRSeg consistently outperforms MetaSeg, a strong recent INR-based segmentation baseline, with an average improvement of 5.6 percentage points in Dice for the in-domain test set and 8.2 percentage points out-of-domain. Overall, our analysis identifies the conditions under which INR-based segmentation is most effective, providing concrete guidance for model selection and future research.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Segmenting parts of brain scans is important for medical care but can be hard because of limited examples and changing scan styles. The authors studied a new type of lightweight model called Implicit Neural Representations (INRs) that can do this segmentation well with fewer parameters. They found INRs work best when limited in size and training data, unlike traditional methods that improve with scale. By combining clues from multiple internal layers, they built a better INR model, HierINRSeg, which performs more accurately on brain MRI scans even across different hospitals.
Open → 2609.31573v1

Object centric angle refinement improves irregular turntable 3D reconstruction

OC-GS: Gaussian Splatting for Irregular Turntable Capture

Abstract: Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores of 21.26, 19.36, and 15.83dB, respectively, exceeding all four evaluated pose-free Gaussian splatting baselines in each condition. Under a shared trainer, refining image-estimated angles improves mean foreground PSNR by 7.88dB over keeping those estimates fixed. An ablation study shows that both image-derived angle initialization and the shared motion model contribute to the improvement. On real captures, OC-GS's refinement increases mean foreground PSNR by 0.70dB. Results show that refining uncertain angles within a shared motion model improves reconstruction from sparse, irregular turntable captures.

Fri 25 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Making 3D models of objects using turntables relies on evenly spaced photos, but uneven spinning and missed images cause problems. The authors developed a method called OC-GS that adjusts each photo’s angle while keeping the camera and rotation settings consistent. This approach helps make better 3D models from fewer and irregular pictures. Testing showed that refining the angles this way gives clearer object images than methods that don’t adjust angles.
Open → 2609.31572v1

Strategically sampled data improves training for complex AI tasks

Strategically Diverse Sampling for Self-Training

Abstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.

Fri 25 SeptComputation and Language
The gist
Training AI models often involves using many examples that look very similar, which doesn’t always help them learn better. The authors studied how showing the AI examples that use very different problem-solving approaches, even if some are incorrect, can improve learning. They created new ways to pick these diverse examples and found that models trained this way perform better on hard problems. Surprisingly, this method can outperform training with much larger models when the data is chosen randomly.
Open → 2609.31571v1

Deep learning improves uncertainty and explanation in option pricing models

Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach

Abstract: Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attributing a single parameter point estimate, it applies local-background Kernel SHAP to a posterior-information functional measuring contraction from the prior to the posterior. This identifies maturity--moneyness regions associated with posterior information gain for individual rHeston parameters. In a simulation study, posterior-predictive intervals provide calibrated or conservative coverage across forward-start, barrier, and realized-variance claims, while point plug-in prices can be materially unreliable for selected contract regimes. Together, the UQ and XAI analyses provide a transparent framework for uncertainty-aware neural calibration and downstream exotic pricing under the specified prior-predictive model.

Fri 25 SeptMachine Learning
The gist
Financial markets often use complex mathematical models to price options, but it's hard to know how uncertain these prices are after using observed market data. The authors develop a new deep learning approach that not only estimates possible values of the model parameters but also quantifies the uncertainty of those estimates. They also introduce a method to explain which parts of the market data influence each parameter estimate. Their approach provides more reliable price estimates and clear insights about uncertainty for different types of financial contracts.
Open → 2609.31570v1

Elementary teachers adapt practices for AI in classroom curriculum

Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum

Abstract: Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate for their students and classroom context. There is little known about how elementary teachers work as they implement conversational AI tools in real classrooms. In this study, we examine three teachers' experiences implementing an AI literacy and English Language Arts (ELA) curriculum built around ToyTalk, a conversational AI toy development platform, over 13 instructional days, a three-week summer camp. Drawing on daily individual reflections, group reflections, and post-camp interviews, we find that teachers' adaptive practices of repair, differentiation, translation, and balancing sit at the intersection of three tensions (technology, learner, and instruction). Teachers' understanding of AI and their role evolved over the camp experiences. From these findings, we contribute design implications and considerations for deploying conversational AI within elementary classrooms.

Fri 25 SeptHuman-Computer InteractionArtificial Intelligence
The gist
AI chat tools are becoming common for kids, and schools want to use them in lessons. This study looks at how three elementary teachers used an AI toy platform during a summer camp and how they changed their teaching to make AI useful and safe for their students. The teachers learned how to fix problems, tailor activities to different learners, explain AI concepts, and balance technology with teaching goals. Their understanding of AI and their teaching role changed as they worked with the AI tool.
Open → 2609.31569v1

DeepEdu improves speed and accuracy of Vietnamese AI tutoring systems

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Abstract: AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.

Fri 25 SeptArtificial Intelligence
The gist
AI tutoring could help students learn better in places like Vietnam, but common approaches have problems with privacy and accuracy on local subjects. The authors created DeepEdu-v1, which works faster and better on Vietnamese curriculum by using a new way to handle long learning conversations and by learning from past interactions without retraining. This reduces data privacy concerns and makes the AI more accurate on complex tasks. Their system cuts response times and improves correctness compared to existing methods.
Open → 2609.31568v1

Mars local computing reduces delays and boosts data from robotic missions

Bandwidth, Latency, and 400 Million Kilometers: The Case for Mars-Local Compute

Abstract: There have been recent proposals for human settlements on Mars in 2030s. Any human activity on Mars must be preceded by extensive robotic exploration. However, Mars exploration is bottlenecked by the low bandwidth, intermittent Mars-Earth link. For example, HiRISE, a high-resolution camera onboard the Martian orbiter MRO imaged less than 3% of Mars over eleven years, even though MRO's low resolution Context Camera had mapped more than 99% of Mars in that time. We present a systems case for shared compute for Mars exploration. Such Mars-local compute, paired with advances in computer vision and AI, can enable large volumes of data to be collected and processed on Mars while sending periodic updates, insights, and selective datasets to Earth. To overcome the lack of surface infrastructure on Mars, we propose a two-tier in-orbit deployment of computational satellites that provides consistent coverage and bandwidth. Our analysis shows that the proposed deployment can start small: one areostationary node makes compute reachable from all active Mars missions, two additional areostationary nodes can extend this coverage to roughly 90% of the planet, while low-Mars-orbit nodes add high-rate surface links and compute capacity where demand grows.

Fri 25 SeptDistributed, Parallel, and Cluster ComputingNetworking and Internet Architecture
The gist
Sending information from Mars to Earth is very slow and often interrupted, which limits how much data robots can send back. The authors suggest placing computers in Mars orbit to process lots of data locally, then only sending important updates to Earth. This setup would include satellites that stay in fixed positions above Mars, covering most of the planet and helping robots explore more effectively. This could make robotic missions much more productive before we send humans there.
Open → 2609.31566v1

Neural network weights optimized to use smaller variable grammar codes

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

Abstract: We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.

Fri 25 SeptMachine Learning
The gist
Deep learning models use many numbers called weights, which can take up a lot of space. The authors show a way to tweak these weights so they can be described by shorter, repeating patterns called grammars, saving memory. They do this by grouping similar weight values together and training the model to work well with those simplified weights. This makes the weight representation smaller but with only a small drop in accuracy.
Open → 2609.31564v1

Multi-agent language models scale differently on two main task types

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

Abstract: Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.

Fri 25 SeptArtificial IntelligenceMultiagent Systems
The gist
Combining multiple AI language models into teams can sometimes improve results, but it depends on the kind of task. The authors studied two task types: disjunctive tasks (where only one team member needs to be right) and compensatory tasks (which depend on averaging answers). They found that for disjunctive tasks, bigger teams rarely improve the most common answer choice, but revising answers with a few peers helps. For compensatory tasks like estimating numbers, averaging answers barely reduces errors because the models tend to make similar mistakes. Mixing different model types helps somewhat for estimation but not for disjunctive tasks.
Open → 2609.31563v1

AI systems coordinate research with resource and risk management

Agentic Economies for Autonomous Scientific Discovery

Abstract: Recent advances in agentic Artificial Intelligence (AI) systems have marked a shift in AI for Science: moving away from the use of individual AI systems for narrow task execution, toward multi-agent systems capable of orchestrating complex, end-to-end research workflows and performing (semi-)autonomous scientific discovery. The development of multi-agent AI-for-science systems has primarily focused on improving the cognitive capabilities of AI systems, specifically by making advanced reasoning and hypothesis generation more reliable. However, focusing only on cognitive capability improvement could ignore appropriate management of resources, a key bottleneck in scientific discovery. Testing and validating scientific hypotheses and experiments is, physically and economically, resource-intensive and resources are limited. This means that to make significant advancements in autonomous scientific discovery, such as improving human-AI co-scientist complementarity or reaching a truly closed-loop automated process, we must pair ongoing improvements in reasoning capabilities with robust resource management. In this paper, we outline an infrastructure for AI resource management by developing the necessary foundations of scientific agent economies, markets, and institutions. The aim of this infrastructure is to empower AI agents and human scientists to effectively (i) collaborate and establish research priorities, (ii) assign credit, (iii) track accountability and liability, and (iv) safeguard against malicious use and information security risks. Finally, we engage with the macro-level societal implications. of (semi-)autonomous scientific discovery to inform the development of governance policies ensuring an equitable distribution of AI-driven discoveries and derivative future technologies.

Fri 25 SeptComputers and Society
The gist
Scientific discovery can be slow and expensive because experiments use real-world resources that are limited. The authors point out that just making AI smarter at thinking and guessing hypotheses is not enough to speed up science. Instead, AI systems need ways to manage resources, share credit, and avoid risks when working together with humans on research projects. They propose building infrastructures like markets and institutions so that AI agents and human scientists can prioritize work, assign credit fairly, and handle security concerns. This approach also considers how society might share the benefits of discoveries made by AI-powered science.
Open → 2609.31562v1

Quantizing neural networks with improved error control using regularization

Generalization behavior of OPTQ and the role of regularization

Abstract: Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the algorithm when a test point is drawn from a fixed distribution. We prove two results. One result relates the generalization error to the error on a calibration dataset comprising independent samples from the same distribution as the test distribution. The other result bounds the generalization error of stochastic OPTQ for all sufficiently nice distributions, regardless of the calibration dataset. In both of these results, the regularization term $λ$ plays an important role. We use insights from these results to make a new recommendation for the choice of $λ$ and see that this choice of $λ$ preforms favorably in experiments when compared to prior recommendations in the literature.

Fri 25 SeptMachine Learning
The gist
Big AI models can be made smaller and faster by turning their many numbers into fewer bits, but this can cause mistakes. The authors study a way called OPTQ that carefully picks how to round these numbers to keep errors low on example data. They prove math results that explain how well this method works on new, unseen data, showing an important role for a penalty term called regularization. Using these insights, they suggest a better value for this penalty and show it works well in tests compared to older suggestions.
Open → 2609.31560v1

Efficient online learning through learned low dimensional parameter tracking

Online Learning via Learned Latent Bayesian Tracking

Abstract: Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or manually designed low-dimensional subspaces. In this work, we identify the absence of a suitable low-dimensional dynamical representation as the core bottleneck in Bayesian filtering-based online learning. Accordingly, we propose Adaptive Update through Representation Adaptation (AURA), a meta-learning framework that learns offline a low-dimensional latent state-space model governing the evolution of optimal model parameters under distribution shift. Online adaptation is then performed via extended Kalman filtering in this learned latent space followed by reconstruction of the full model parameters through a learned lifting map, enabling efficient single-step online adaptation while preserving model expressiveness. Evaluated on online adaptation of neural wireless receivers under time-varying channels and on non-stationary image classification, AURA shows substantial improvements in adaptation speed, accuracy, and computational efficiency over existing online learning and Bayesian filtering baselines, demonstrating that an adaptation-aware latent geometry is beneficial for effective Bayesian online learning in high-dimensional models.

Fri 25 SeptMachine Learning
The gist
Online learning means updating a model quickly as new information comes in, but it is very hard to do this with big deep learning models because they have too many parts to adjust at once. The authors found that the problem is not having a simple enough way to represent how the best model parameters change over time. They created a system called AURA that learns a smaller, simpler map of these changes ahead of time, so it can update the full model quickly when new data arrives. This helps models adapt faster and with less computation in changing environments like wireless signals or shifting image categories.
Open → 2609.31559v1

CLIPGuard defends image AI from hidden embedding space backdoors

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

Abstract: Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git

Fri 25 SeptComputer Vision and Pattern RecognitionCryptography and Security
The gist
Some bad actors hide secret triggers in images that fool AI models like CLIP, making them behave wrongly. Existing defenses often need to see inside the model or lots of clean data, which is not always possible. The authors designed CLIPGuard, a tool that treats the model as a black box and finds suspicious parts of an image to fix them without hurting the good parts. Their tests show CLIPGuard can stop these hidden backdoors effectively while keeping the AI’s normal accuracy high.
Open → 2609.31558v1

Mexican spanish video dataset advances hate speech detection research

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

Abstract: Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.

Fri 25 SeptComputation and LanguageComputer Vision and Pattern Recognition
The gist
Detecting hate speech online is important but tricky because it depends on subtle cultural and language clues. The authors created MexHat, a collection of about 1,000 video clips in Mexican Spanish labeled to show whether they contain no negativity, offensive language, or hate speech. The dataset also breaks down hate speech into detailed categories. This resource helps build and test tools that understand hate speech better in Mexican Spanish videos.
Open → 2609.31553v1

Large language models can be tricked to waste more computing power

FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation

Abstract: As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting their deployment-wide impact when benign traffic dominates. In this work, we uncover a previously overlooked token-level attack surface arising from the many-to-one mapping from token sequences to decoded text. Although standard LLMs predominantly generate the canonical token sequences induced by their tokenizers, the same text can also be represented by substantially longer non-canonical sequences. This representational flexibility exposes a new avenue for resource-consumption attacks: an attacker can train the model to favor such sequences, systematically increasing the number of autoregressive decoding steps without a proportional increase in visible response length. However, we empirically find that directly maximizing token fragmentation substantially degrades model utility, producing conspicuous answer-quality failures that undermine attack stealthiness. To address this challenge, we propose FragToken, a training-time framework that combines source-model self-distillation, capacity-aware filtering and budgeting, and BPE-Aligned Merging to induce fragmented generation under ordinary prompts while largely preserving model utility. We evaluate FragToken on four LLMs across three benchmarks. Across the four models, FragToken achieves a three-benchmark average token inflation ratio (TIR) ranging from 1.99 to 2.46, while causing only minor degradation in model utility. Our work reveals a covert LLM supply-chain threat that increases inference cost without requiring large volumes of attack requests while largely preserving utility.

Fri 25 SeptCryptography and Security
The gist
Large language models (LLMs) are costly to run, so attackers try to make them use more resources. The authors found a hidden trick where the model can produce the same visible text but using many more internal steps. This means the model works harder without showing longer outputs, making attacks harder to detect. They created FragToken, a method to train models to do this token trick while still giving useful answers, proving this hidden attack is real and effective.
Open → 2609.31552v1

Efficient serving system balances GPU use for multimodal language models

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Abstract: Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.

Fri 25 SeptDistributed, Parallel, and Cluster ComputingMachine LearningPerformance
The gist
Multimodal large language models process images, video, or audio along with text, requiring an extra Encode step that turns media into embeddings. Existing serving systems struggle to allocate GPU resources well among Encode, Prefill, and Decode steps, leading to wasted capacity. The authors propose EAServe, a system that controls when and how work flows through these three steps and allocates GPUs dynamically to keep all stages busy. This approach improves throughput and GPU utilization compared to prior methods under the same delay goals. EAServe was tested on models handling image, video, and audio inputs, showing faster and more balanced performance.
Open → 2609.31551v1

Two methods control false alerts when screening AI text inside documents

Two Conformal Constructions for Adaptive Within-Document AI-Text Screening

Abstract: We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.

Fri 25 SeptComputation and Language
The gist
Detecting text generated by AI inside documents can raise false alarms, wasting effort. The authors propose two ways to carefully decide when to stop checking parts of a document to control false alerts, even if tokens inside the document depend on each other. They prove that their methods keep false alarms under control while allowing early stopping, but testing their accuracy and speed in practice is still needed.
Open → 2609.31547v1

Beatgraph models infant heartbeats directly to improve ecg analysis

BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment

Abstract: Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats directly rather than recover them from arbitrary patches. We propose BeatGraph, which makes the heartbeat its unit of representation, modeling each 30-second window as a graph of beats. A shared beat encoder embeds each heartbeat from its waveform and inter-beat intervals, a Transformer with positional encoding orders the beats in time, and residual graph attention layers relate every beat to every other before attention pooling yields a window embedding. We pretrain BeatGraph on our new corpus of unlabeled infant recordings by predicting masked-beat embeddings, then fine-tune it for each task. One backbone supports sleep-wake detection, infant-state classification, activity-source identification (infant- or caregiver-initiated movement), and affect recognition, improving macro-F1 over the strongest baseline on each task by 0.076 to 0.158. It also transfers across age groups, reaching 0.892 AUROC on the ZZU-pECG pediatric benchmark (ages 0 to 14), within 0.001 of the best published self-supervised ECG model, and matching that model under linear evaluation on the adult PTB-XL benchmark despite infant-only pretraining. Finally, to our knowledge, we release the first public infant ECG corpus collected in homes, classrooms, and laboratory settings with state and affect labels. It contains 3,408 hours of single-channel ECG from 143 infants aged 3 to 11 months, with unlabeled pretraining data, benchmark tasks, and subject-level splits.

Fri 25 SeptMachine Learning
The gist
Infant heartbeats are faster and different from adult heartbeats, but many ECG models ignore these differences by cutting the signal into random pieces. The authors created BeatGraph, which looks at each individual heartbeat and how they connect over time to better understand baby ECG signals. They trained this new model on a large collection of baby ECG data recorded at home and other everyday places. BeatGraph works well to detect sleep, activity, and emotions in infants, and it also performs strongly when tested on older children and adults.
Open → 2609.31546v1