AI research / 2026-10-07arXiv:2610.08789v1 Announce Type: cross Abstract: Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08785v1 Announce Type: new Abstract: Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08773v1 Announce Type: cross Abstract: Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08764v1 Announce Type: cross Abstract: We develop the first feedback design for rapid stabilization of the Kuramoto--Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron and L\"u…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08750v1 Announce Type: new Abstract: Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08745v1 Announce Type: new Abstract: We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08743v1 Announce Type: new Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08740v1 Announce Type: new Abstract: Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08738v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08735v1 Announce Type: new Abstract: In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08722v1 Announce Type: cross Abstract: Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08718v1 Announce Type: cross Abstract: Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself:…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08717v1 Announce Type: cross Abstract: We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08715v1 Announce Type: cross Abstract: The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08694v1 Announce Type: new Abstract: Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08689v1 Announce Type: new Abstract: Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08680v1 Announce Type: new Abstract: Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08678v1 Announce Type: cross Abstract: Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08677v1 Announce Type: new Abstract: Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08670v1 Announce Type: new Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08669v1 Announce Type: new Abstract: On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08652v1 Announce Type: cross Abstract: Diffusion models are increasingly used as surrogates for expensive simulators in weather prediction, molecular dynamics, and materials design. In these models, computing the probability $p_0[E]$ of an event $E$ is difficult,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08647v1 Announce Type: cross Abstract: LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08626v1 Announce Type: cross Abstract: Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08624v1 Announce Type: new Abstract: We propose a method for choosing the shared memory parameter $\beta_1=\beta_2=\beta$ in Adam from a short pilot training. The selected $\beta$ remains fixed during the subsequent full training. A local model of Adam's normalized…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08593v1 Announce Type: new Abstract: Traditional approaches for automated bug detection in video games, such as manual testing, can be beneficial for the improvement of quality assurance, but they can be expensive and time-consuming. The scarce number of tools…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08592v1 Announce Type: new Abstract: CNet is a C++/CUDA framework for building and training deep complex-valued neural networks (CVNNs) and, more generally, for optimizing complex-valued functions by gradient descent with Wirtinger (CR-calculus) derivatives. It takes…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08578v1 Announce Type: new Abstract: Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08577v1 Announce Type: new Abstract: While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08574v1 Announce Type: cross Abstract: Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08571v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08570v1 Announce Type: new Abstract: Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08565v1 Announce Type: new Abstract: This article is a geometric rediscovery of the singular value decomposition, with a further claim: the construction it builds is the machinery behind much of machine learning. The same argument that answers an idle question about…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08564v1 Announce Type: new Abstract: Tabular foundation models (TFMs) can classify the nodes of a graph without training on it, by reading node and neighborhood features as table rows next to labeled context rows. Work in this line reports predictive performance, not…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08561v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08560v1 Announce Type: cross Abstract: Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08559v1 Announce Type: cross Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08554v1 Announce Type: cross Abstract: While HCI increasingly examines AI-safety for youth, the literature lacks a comprehensive view of what risks have been identified, how they are addressed, and whether proposed protections work in-practice. We systematically…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08553v1 Announce Type: new Abstract: Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08552v1 Announce Type: cross Abstract: Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08540v1 Announce Type: cross Abstract: Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08538v1 Announce Type: new Abstract: Probabilistic load forecasting has been widely studied for power-system operation and planning, but customer- and transformer-level forecasting introduces a distinct scalability challenge. At these levels, load uncertainty is…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08537v1 Announce Type: new Abstract: In the field of Explainable AI (XAI), counterfactual (CF) explanations interpret a model's decision by suggesting the changes to the input that would lead to a more favourable outcome. To be useful in practice, such an explanation…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08534v1 Announce Type: new Abstract: Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08533v1 Announce Type: cross Abstract: Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08527v1 Announce Type: new Abstract: Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA)…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08502v1 Announce Type: cross Abstract: Proactive power management systems reduce processor dynamic power through runtime power prediction and power-aware scheduling. Accurate, stable and low-overhead digital on-chip power meters (OPMs) are crucial for improving the…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08495v1 Announce Type: cross Abstract: Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08479v1 Announce Type: new Abstract: Few-shot meta-learning traditionally formulates task adaptation either as analytical gradient descent through unrolled computational graphs or as metric-based distance comparisons over flattened 1D fea- ture vectors, which either…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08475v1 Announce Type: new Abstract: In many-query scenarios, data-driven surrogate models provide an efficient alternative to high-fidelity solvers for simulating physical systems governed by Partial Differential Equations (PDEs). In this context, the Latent Dynamics…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08463v1 Announce Type: cross Abstract: Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08457v1 Announce Type: new Abstract: Integrated circuit design involves multiple design stages: logic synthesis, floorplanning, placement, and routing, with each stage taking hours to weeks to complete. Discovering timing violations late in this flow forces costly…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08452v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08420v1 Announce Type: new Abstract: We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08413v1 Announce Type: cross Abstract: Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08403v1 Announce Type: new Abstract: Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08402v1 Announce Type: new Abstract: Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome,…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08400v1 Announce Type: new Abstract: Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to generalize across downstream tasks. While deep learning has driven considerable…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08384v1 Announce Type: new Abstract: In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.08368v1 Announce Type: new Abstract: Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives. To navigate this multidimensional space, Corbion and…
arXiv · cs.LG ↗AI research / 2026-10-07arXiv:2610.06852v1 Announce Type: cross Abstract: Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06851v1 Announce Type: cross Abstract: In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06846v1 Announce Type: new Abstract: Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06844v1 Announce Type: cross Abstract: Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06843v1 Announce Type: cross Abstract: LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06831v1 Announce Type: cross Abstract: Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight,…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06830v1 Announce Type: cross Abstract: Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06829v1 Announce Type: cross Abstract: Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06824v2 Announce Type: new Abstract: We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06823v1 Announce Type: cross Abstract: Large longitudinal cohorts often contain wrist accelerometry without optical heart-rate sensing, motivating recovery of cardiac information from motion signals already collected during sleep. We present SeqSmoother, a…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06817v1 Announce Type: cross Abstract: We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06814v2 Announce Type: cross Abstract: World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06804v1 Announce Type: cross Abstract: A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06801v1 Announce Type: cross Abstract: Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06790v1 Announce Type: new Abstract: The unprecedented computational scale of modern artificial intelligence depends on complex, multi-billion-transistor Systems-on-Chip, yet the workflows that verify these chips remain stubbornly manual. Although Large Language…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06778v1 Announce Type: cross Abstract: While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06765v1 Announce Type: new Abstract: Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06751v1 Announce Type: cross Abstract: Matrix completion underlies problems from tabular imputation to causal inference, yet existing tabular foundation models treat it as entry-by-entry prediction, repeating context for every target and discarding the matrix's…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06750v1 Announce Type: cross Abstract: Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06748v1 Announce Type: cross Abstract: In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06725v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06679v1 Announce Type: cross Abstract: Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06673v1 Announce Type: new Abstract: Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06672v1 Announce Type: cross Abstract: Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms:…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06668v1 Announce Type: new Abstract: Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06652v1 Announce Type: new Abstract: Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06643v1 Announce Type: cross Abstract: Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06636v1 Announce Type: cross Abstract: Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public"…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06623v1 Announce Type: cross Abstract: Building on our earlier program-evolution workflow guided by large language models (LLMs), we study weight-five bivariate bicycle (BB) and perturbed bivariate bicycle (PBB) codes. The resulting catalogue contains 1,142 distinct…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06621v1 Announce Type: cross Abstract: Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06614v1 Announce Type: new Abstract: As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06603v1 Announce Type: cross Abstract: Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06598v1 Announce Type: cross Abstract: Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06597v1 Announce Type: new Abstract: LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06591v1 Announce Type: new Abstract: How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06587v1 Announce Type: cross Abstract: AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06584v1 Announce Type: cross Abstract: The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06582v1 Announce Type: new Abstract: World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06578v1 Announce Type: new Abstract: While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06571v1 Announce Type: cross Abstract: Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06563v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06553v1 Announce Type: new Abstract: Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06549v1 Announce Type: new Abstract: Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded knowledge, conclusions, and recommendations. We present the CKG Clinical Explanation Engine, a downstream layer for a frozen,…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06535v1 Announce Type: cross Abstract: Magic-state distillation is a major resource cost in fault-tolerant quantum computing. The cost of a magic-state factory depends strongly on its failure rate, which grows with the number of input magic states. Although…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06522v1 Announce Type: cross Abstract: A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06517v1 Announce Type: cross Abstract: IoT encompasses diverse physical entities, from smart home devices to autonomous vehicles, creating a complex environment with heterogeneous security models. This heterogeneity makes IoT sub-systems vulnerable to various network…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06514v1 Announce Type: new Abstract: The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06511v1 Announce Type: cross Abstract: To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06503v1 Announce Type: cross Abstract: Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06496v1 Announce Type: new Abstract: When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06494v1 Announce Type: cross Abstract: Multi-structure segmentation of the uterus is important for computer-assisted screening, diagnosis, and treatment planning of uterine diseases, where ultrasound and MRI provide complementary clinical information. However,…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06491v1 Announce Type: new Abstract: Multi-view learning jointly exploits multiple complementary representations of the same data and has become increasingly important in machine learning. However, the Python ecosystem lacks actively maintained, unified tooling for…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06489v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06469v1 Announce Type: cross Abstract: Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06468v1 Announce Type: cross Abstract: Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06464v1 Announce Type: cross Abstract: As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06454v1 Announce Type: new Abstract: The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06453v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) achieve strong generalization by learning to reconstruct or forecast broad temporal patterns from large-scale time series during pre-training. Yet this strength can become a weakness for…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06452v1 Announce Type: new Abstract: VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that…
arXiv · cs.AI ↗AI research / 2026-10-07arXiv:2610.06446v1 Announce Type: cross Abstract: Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors…
arXiv · cs.AI ↗