ArXiV ML/AI/CV papers summary
Theme 1: Resource-Aware Inference and Efficient Deployment
The “memory wall” and thermal constraints of edge devices have necessitated a shift from monolithic, cloud-centric models to intelligent, multi-tier systems. By treating physical constraints—such as thermal headroom—as state variables, we can optimize performance without sacrificing hardware integrity.
- Thermal-Aware Routing: HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference
- Efficient Architecture Search: ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
- KV Cache Optimization: The KV Cache Is the New Memory Wall and CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
- Multimodal Efficiency: SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages and I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
Theme 2: Agentic Reasoning, Tool-Use, and Executive Control
We are witnessing the evolution of AI from passive chatbots to autonomous agents capable of planning, tool-use, and long-horizon reasoning. This transition requires governance architectures that prevent “executive-control failure”—where agents waste resources on redundant refinements—and frameworks that enable effective tool selection.
- Executive Control: LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents
- Reasoning and Deferral: Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents and Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
- Tool-Use and Skill Evolution: ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning, MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning, CODESKILL: Learning Self-Evolving Skills for Coding Agents, and SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
- Practical Applications: REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles, Epstein Files Engine: Agentic Search for Investigative Journalism, and WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery
Theme 3: Mechanistic Interpretability, Causal Steering, and Learning Dynamics
To move beyond “black-box” models, we must understand the geometry of internal states. This allows for targeted steering and a more formal science of neural representation, supported by rigorous optimization theories that explain why specific training techniques succeed.
- Steering and Control: Steering Interference Reflects the Model’s Defaults, Not the Behavior Directions and LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
- Causal Retention: Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation
- Algebraic Frameworks: Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
- Optimization Dynamics: Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness, Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods, and Toward Understanding Momentum Acceleration in River-Valley Loss Landscape
Theme 4: Physics-Informed AI and Scientific Discovery
Integrating physical laws into neural networks allows for more reliable simulations and scientific discovery. By embedding physical equations directly into the training process, we create models that are not only accurate but also physically consistent.
- Physics-Informed Neural Networks (PINNs): Staged Depth Training: A Representation Curriculum for PINNs and Gradient Surgery for Physics-Informed Neural Networks
- Scientific Emulation: ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets and Latent Generative Solvers for Generalizable Long-Term Physics Simulation
- Molecular and Domain-Specific Science: Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark, Neural Parameter Estimation of RC Thermal Building Models for Model Predictive Control, and Insurance Reserve Intelligence Platform
- Embodied Scientific Agents: SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents and SUN: Agentic Robot Policy Learning with Persistent Task Programs
Theme 5: Safety, Alignment, and Temporal Reliability
As AI systems interact with the real world, safety must be integrated into the architecture rather than treated as a post-hoc filter. This includes addressing temporal failure modes, where models rely on stale information, and ensuring robust behavior under distribution shifts.
- Safety and Alignment: The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists, ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?, and Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
- Temporal Reliability: Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers, Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering, and Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
- Robustness and Bias: Conformal Prediction under Exponential-Tilt Joint Shift, Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift, Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning, RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models, and Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
Theme 6: Embodied AI and Advanced Evaluation
The frontier of robotics involves “World Action Models” that simulate consequences before execution. Simultaneously, the field is moving toward dynamic, process-oriented evaluation to replace saturated static benchmarks.
- World Models: MA-WAM: Multi-Agent World-Action Model for Test-Time Planning, AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control, Keep the Future, Drop the Rollout: RIFT for World Action Models, and SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
- Evaluation and Benchmarking: Game Arena: Strategic LLM Evaluation in Competitive Environments, Stepwise Intrinsic Rewards for Reasoning in Large Language Models, CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating, TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding, and STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence