ArXiV ML/AI/CV papers summary
Theme 1: Agentic Reasoning, Planning, and Reliability
The frontier of AI is shifting from simple “next-token” prediction to autonomous, multi-step agentic systems. The core challenge lies in moving beyond statistical mimicry toward physically grounded, long-horizon reasoning that remains reliable under real-world constraints.
- Reliability & Verification: Sentry: Learning to Recover from LLM Agent Failures at Test Time and DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents introduce failure-management and modular gating to validate actions. Safety is further addressed by A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns, which highlights the need for multi-layer defenses against long-term constraint violations.
- Planning & Self-Improvement: Research is moving toward recursive optimization and self-honing. Recursive Agent Optimization, HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents, and ASH: Agents that Self-Hone in Long-Horizon Worlds allow agents to spawn sub-tasks and learn from failure-relevant trajectories. MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation and Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners emphasize the necessity of decoupling high-level reasoning from low-level control.
- Benchmarking: OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine, SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation, and Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks warn that current benchmarks often fail to capture real-world complexity or suffer from representation-dependent artifacts.
Theme 2: Embodied Intelligence and Physics-Informed AI
We are witnessing a convergence where AI architectures are no longer “black boxes” but systems that respect the laws of physics, enabling digital twins and sophisticated robotic manipulation.
- Robotic Manipulation & 3D Awareness: PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation and MoSE3: Learning World-Space SE(3) at Every Pixel integrate 3D spatial awareness to forecast physical evolution. World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation, GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation, and EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos bridge the gap between high-level intent and physical interaction.
- Scientific Discovery & Simulation: The AI Theorist reveals excitonic structure in $\alpha$-RuCl$_3$, MolWorld: Molecule World Models for Actionable Molecular Optimization, and Coupled reaction and diffusion governing interface evolution in solid-state batteries demonstrate AI’s role in autonomous scientific discovery. Architectures like S$^{2}$-PINN: Stochastic Separable Physics-Informed Neural Networks, Dirac-Interconnected Neural Elements: Discovering Modularity in Physical Systems Without Reduction, and Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals respect physical constraints while reducing computational costs.
- Generative Plausibility: ProAR: Learning Prospective Reasoning with Autoregressive Video Models, FlowHMR: Physically Plausible Motion Capture from Video, and I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry ensure that generated content is physically consistent and simulation-ready.
Theme 3: Efficiency, Inference, and Memory Management
As models scale, the “KV-cache bottleneck” and computational overhead have become primary constraints. Research is shifting toward hardware-aware optimization and inference-time scaling.
- KV-Cache & Memory: SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention, Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention, iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD, and CORE: COverage CAlibration and Evicted-Mass REdistribution for KV Cache provide sophisticated methods to manage memory without sacrificing reasoning quality.
- Hardware & Inference Acceleration: Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts, Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights, BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration, and Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning optimize the mechanics of model execution.
- Adaptive Computation: SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching, Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection, FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution, and LensVLM: Selective Context Expansion for Compressed Visual Representation of Text demonstrate that efficiency is achieved by allocating resources only where they are needed most.
Theme 4: Trust, Safety, and Causal Reasoning
The field is maturing toward verifiable AI, where models must be auditable, robust to adversarial attacks, and grounded in causal reality.
- Security & Alignment: MLCommons Jailbreak Benchmark v1.0, Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models, Securing Computer-Use Agents Against Branch Steering Attacks, and AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents address the expanding attack surface of agentic systems. Suan: Rectifying Direct Preference Safety Alignment in Large Language Models and Mitigating Social Sycophancy via Pluralistic Preference Optimization tackle alignment and sycophancy.
- Faithfulness & Causal Discovery: LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs and Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content provide tools to audit hallucinations. Counterfactual Predictions in Scientific Emulators Without Controlled Experiments and Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery push the boundaries of what-if reasoning.
- Statistical Reliability: Pragmatic DML with AI-Learned Representations, DeepGOF: Where Does a Logistic Risk Model Fail?, and EDGE: A grouped calibration test for logistic regression integrate deep learning into rigorous statistical frameworks, ensuring models are as reliable as traditional scientific methods.
Theme 5: Multimodal Reasoning and Adaptation
The integration of diverse sensory inputs is moving toward “omni-modal” intelligence that maintains consistency across modalities.
- Multimodal Foundation Models: OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection, SCOPE-4D: Endoscopic 4D Geometry Foundation Models, and EXAM2: Extending Audio Understanding in Multilingual and Multimodal Analysis adapt foundation models to specialized domains.
- Hallucination & Uncertainty: OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination, CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models, and DICE: Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning provide mechanistic insights into failure modes.
- Efficient Adaptation: Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs, Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization, and LightLoc++: Sensor-Robust Representation Learning for Efficient Outdoor LiDAR Localization demonstrate that specialized pathways and robust backbones are key to efficient deployment.