ArXiV ML/AI/CV papers summary
Theme 1: Agentic Reasoning, Orchestration, and Reliability
The field is undergoing a fundamental transition from passive “chat” models to autonomous agents capable of long-horizon planning, tool use, and self-correction. As these agents gain execution authority, the focus has shifted from simple capability to structural reliability and governance.
- Architectural Patterns: To manage complexity, researchers are moving toward modular, orchestrated frameworks. A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics and ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence demonstrate how specialized agents can be orchestrated to solve multi-step problems. Supporting this, MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents and BusMA: A Bus Communication Substrate for Multi-Agent Systems provide the “plumbing” for inter-agent communication and skill sharing.
- Reliability & Verification: Simple “generate-test-revise” loops are insufficient for high-stakes environments. Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair and Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management emphasize the need for “decision assurance” and “recoverability.” Furthermore, Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit and Auditable Agents argue that accountability requires decoupling the agent’s persona from its audited execution path.
- Safety & Security: Protecting against indirect prompt injection (IPI) is critical. ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents and DualView: Preventing Indirect Prompt Injection in Personal AI Agents propose pre-execution audits and data-view separation to prevent malicious hijacking. HazardAuditor: From Executable Threats to Safer Computer-Use Agents and Safety Signals to Verify NetOps Agents with Action-Level Granularity advocate for execution-grounded monitoring.
Theme 2: Efficient Inference, Optimization, and Memory
As models scale, the bottleneck shifts to latency, memory management, and the cost of “test-time compute.”
- Inference Efficiency: LLM Inference in a Flash! and Breaking the 1.58-bit Barrier for Ternary LLMs push quantization limits, while Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching and Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction optimize the KV cache. T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing and MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving use dynamic routing to achieve speedups without sacrificing accuracy.
- Memory Architectures: Moving beyond “append-only” streams, The Immutable Past: Formalizing State Mutability and Conflict Resolution in Mutable RAG and Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents propose active memory systems that reconsolidate and compress information, mimicking cognitive neuroscience to maintain long-term coherence.
- Mathematical Optimization: Learning Choice Model Trees for Feature-Based Multi-Product Pricing: Exact Optimization and Field Evidence and A proximal augmented Lagrangian method for nonconvex optimization with equality and inequality constraints move away from greedy heuristics toward exact, mathematically grounded optimization methods.
Theme 3: Physics-Informed and Scientific Machine Learning
Integrating physical laws into neural architectures ensures reliability and data efficiency in scientific domains.
- Surrogate Modeling: Neural Field Ensembles for Aerodynamic Surface Prediction: Winning Solution to the ONERA CRM Wall Distribution 2025 Challenge and A panoramic aerodynamic performance prediction method for turbomachinery cascades using transformer-enhanced neural operator demonstrate that neural operators can act as high-fidelity simulators.
- Constraint Satisfaction: Scaling Laws for Physics-Aware ACOPF Surrogate Learning and Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling show that embedding domain constraints directly into training objectives is essential for reliable deployment.
- Embodied AI: ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies and Proprioception-Anchored Cross-Modal Pretraining for Zero-Shot Sim-to-Real Contact-Rich Assembly ensure that physical agents remain grounded and safe through hard constraint enforcement and consistent proprioceptive anchoring.
Theme 4: Statistical Inference, Causal Modeling, and Interpretability
Understanding the “why” behind model behavior is becoming as important as the “what.”
- Causal Inference: Causal Path Analysis from Perturbational and Population-Scale Single-Cell Data with Multiscale Confounding and Measurement Error and Covariate Selection for Doubly Robust Double/debiased Machine Learning Estimators for Causal Inference provide rigorous frameworks for untangling complex, high-dimensional causal relationships.
- Geometric & Statistical Insights: Large Language Models Develop Belief State Geometry In-Context and Characterizing Heterogeneous Rates in Finite Mixture Estimation via Partial Optimal Transport explore the internal geometry of models, while Statistical Inference for Score Decompositions allows for the granular auditing of model performance across miscalibration and uncertainty.
- Interpretability: ICON Decomposition: Auditing deep neural networks for shortcuts by decomposing layer-wise representations using concepts and A unified framework for global and local interpretability using adaptive derivative-ordered random explanation offer new ways to audit models for spurious correlations, moving beyond simple feature attribution.
Theme 5: Alignment, Bias, and Human-AI Interaction
Ensuring models remain helpful, honest, and aligned with human intent requires moving beyond static safety filters.
- Alignment Challenges: Playing Devil’s Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy and Alignment Whack-a-Mole: Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models highlight the fragility of current alignment techniques, where safety is often a “veneer” that can be bypassed by fine-tuning.
- Auditing Bias: Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models and FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation reveal that audit methodologies themselves can introduce bias, necessitating more standardized, context-aware evaluation frameworks.
- Human-in-the-Loop: Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation and Personalizing Personal Health Interfaces: Co-Design with Generative AI emphasize that technical safety is insufficient without inclusive deliberation and user-centric design.