ArXiV ML/AI/CV papers summary
Theme 1: Agentic Reasoning, Tool-Use, and Planning
The paradigm of AI is shifting from monolithic, passive text generation to active, agentic workflows. These systems are designed to plan, utilize external tools, and verify their own outputs, moving toward a future where AI acts as a reliable collaborator rather than just a chatbot.
- Reliability & Verification: To move beyond the “black-box” nature of LLMs, researchers are implementing rigorous verification layers. CURA: Certified Runtime Alarms for Computer-Use Agents and The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs provide mechanisms to detect and control errors in real-time. VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning improves credit assignment in long-horizon tasks, while PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation and MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation use formal verification (LTL/FSM) to ensure safety before physical execution.
- Architectural Efficiency: Managing context over long horizons is a primary bottleneck. TokenPilot: Cache-Efficient Context Management for LLM Agents and String: An Agentic OS Where Every App Is a Markdown File offer structural solutions for context accumulation. HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees optimizes training through prefix-sharing.
- Governance & Control: As agents gain autonomy, security must be architectural, not just prompt-based. Agentao: A Policy-Governed Runtime Harness for Embeddable Tool-Using LLM Agents and Prompts Don’t Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control advocate for proxy-enforced access control to prevent unauthorized tool use.
- Benchmarking: Current benchmarks often fail to capture real-world complexity. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests, LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis, and TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development highlight the gap between synthetic tests and expert-level human performance.
Theme 2: Embodied AI and Physical Reasoning
Intelligence is increasingly grounded in the physical world, requiring models to understand 3D space, temporal dynamics, and causal sequences.
- Robotic Manipulation & Navigation: PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models and RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation focus on anticipating future dynamics and learning from failure. Action- and Language-Conditioned Video Assessment for Embodied Control introduces ALVA to evaluate task progress through a more nuanced, human-like understanding of visual transitions.
- 3D Understanding & World Modeling: SpatialCrafter: Single Image World Modeling with Generative 3D Proxies uses 3D proxies to prevent hallucinations in video generation, while InstructMesh: Selective Refinement of Generative 3D Models for Fabrication and Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning bridge the gap between visual plausibility and geometric/executable accuracy.
- Temporal Awareness: Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning reveals that models struggle with logical sequencing, a finding supported by AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction, which factorizes motion and appearance to improve temporal prediction.
Theme 3: Optimization, Efficiency, and Architecture
The “optimizer” has evolved into a system-level component, and model architectures are being redesigned for speed, modularity, and hardware co-design.
- Optimization Paradigms: Blog: Survey of Optimizers highlights the rise of matrix-aware methods like Muon. When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging and Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining address catastrophic forgetting and ill-conditioned loss landscapes.
- Efficiency & Quantization: DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization, A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint, and HyQuant: Hybrid-Precision Quantization for LLM Attention demonstrate that intelligent, layer-wise bit allocation is critical for edge deployment.
- Hardware & Inference: Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI showcases autonomous hardware-software co-design. Trajectory-Level Speculative Decoding for Diffusion Language Models and Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration provide significant speedups for inference and training.
Theme 4: Scientific Machine Learning and Statistical Rigor
Machine learning is becoming a powerful tool for scientific discovery, provided it is anchored in physical laws and statistical integrity.
- Physics-Informed Models: Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics, Hypergraph Adaptive Wavelet Operators for Parametric PDEs, QGPINNs: A Physics-Informed Neural Network Framework for Nonlocal Differential Equations on Quantum Graphs, and WINO: A Weak-Form Physics Informed Neural Operator for Hyperelasticity on Variable Domains demonstrate how incorporating geometry and physical constraints leads to robust simulations. Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields and MAGE: Multimodal Agentic Framework for Discovering Governing PDEs automate the discovery of governing equations.
- Statistical Foundations: A Deep Zero-Inflated Model of North Atlantic Right Whale Presence, Autotune: fast, accurate, and automatic tuning parameter selection for Lasso, and On efficiency gains via augmenting a tiny sample with a massive auxiliary sample emphasize that deep learning must be supported by rigorous statistical methods. The role of parameter Jacobians in the stability of network outputs and Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions provide the theoretical stability required for reliable learning.
Theme 5: Trust, Safety, and Alignment
As AI systems become more autonomous and capable, ensuring they remain aligned with human values and robust against adversarial threats is a critical research frontier.
- Safety & Robustness: REPLICANT: Learning Policies for Evading and Hardening Malware Detectors, SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces, and TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models expose new attack surfaces. The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior warns that current safety evaluations are often too stochastic to be reliable.
- Interpretability & Alignment: CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation and Bit-Level Triangular Content-Aware Permutation for Fragile Image Watermarking push for transparency and forensic integrity. Large Reasoning Models Learn Better Alignment from Flawed Thinking, Performative Privacy: When Differential Privacy Maximizes Utility, and JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification offer frameworks for better alignment and factuality checking.
- Sycophancy & Bias: The Effect of Emotional Context on Large Language Models’ Endorsement of Premature Decisions and The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models highlight the persistent challenge of models prioritizing user-pleasing over objective truth.