AI research / 2026-10-07arXiv:2610.08789v1 Announce Type: cross Abstract: Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08785v1 Announce Type: new Abstract: Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08773v1 Announce Type: cross Abstract: Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08764v1 Announce Type: cross Abstract: We develop the first feedback design for rapid stabilization of the Kuramoto--Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron and L\"u…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08750v1 Announce Type: new Abstract: Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08745v1 Announce Type: new Abstract: We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08743v1 Announce Type: new Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08740v1 Announce Type: new Abstract: Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08738v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08735v1 Announce Type: new Abstract: In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08722v1 Announce Type: cross Abstract: Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08718v1 Announce Type: cross Abstract: Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself:…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08717v1 Announce Type: cross Abstract: We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08715v1 Announce Type: cross Abstract: The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08694v1 Announce Type: new Abstract: Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08689v1 Announce Type: new Abstract: Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08680v1 Announce Type: new Abstract: Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08678v1 Announce Type: cross Abstract: Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08677v1 Announce Type: new Abstract: Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08670v1 Announce Type: new Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08669v1 Announce Type: new Abstract: On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08652v1 Announce Type: cross Abstract: Diffusion models are increasingly used as surrogates for expensive simulators in weather prediction, molecular dynamics, and materials design. In these models, computing the probability $p_0[E]$ of an event $E$ is difficult,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08647v1 Announce Type: cross Abstract: LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08626v1 Announce Type: cross Abstract: Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature…
arXiv · cs.LG ↗