ArXiV ML/AI/CV papers summary
Theme 1: Scientific Machine Learning & Physics-Informed Discovery
We are witnessing a profound shift from “black-box” pattern matching to models that respect the fundamental laws of the universe. By embedding physical constraints directly into neural architectures, we are moving toward systems that can reliably simulate dynamical systems and accelerate scientific discovery.
- Stability and Physics: Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers and RAPTOR: RAndom-projection Physics-informed Transient sOlveR emphasize that stable long-horizon evolution requires more than just reconstruction; it demands noise injection and physics-informed constraints. Similarly, Physics and Data Driven Transformer-Mamba Framework for Flow Field and Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography warn that low image error is a poor proxy for physical truth, necessitating cross-seed instability diagnostics.
- Scientific Autonomy: AI is becoming a partner in the deductive process. Learning to Discover Interesting Mathematics, A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension, and Long-horizon autoformalization of a core theorem underlying MIP* = RE demonstrate that agents can now autonomously generate conjectures and complete machine-checked proofs. In the physical sciences, Neural-Network Solutions to Real-Space Charge Density and Generalization, iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model, and SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback showcase how recursive AI development and operator graphs are accelerating materials science and industrial coding.
- Clinical Reliability: Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark and An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer provide the infrastructure for high-stakes medical environments, prioritizing auditability and physician-in-the-loop validation.
Theme 2: Agentic Reasoning, Reliability, and Governance
The frontier of AI is no longer just text generation; it is the creation of persistent, tool-using agents. As these systems move into the real world, the focus has shifted from “capability” to “accountability.”
- Reliability and Verification: Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents and GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI highlight the need for rigorous tool contracts and decoupled planning. RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers? and RRSI: Regularized Recursive Self-Improvement of Agent Harnesses address the challenges of self-improvement and reproducibility.
- Governance and Security: We must treat governance as an executable contract rather than a suggestion. Who Holds the Pen? Let Specifications, Not Agents, Sign Off and The Last Human Gate: Forward Deployed Engineering for Governance Automation argue for separating the actor from the authority. AgentKernel: The Trust-Native Agentic Operating System, Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution, and Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture propose moving security into the substrate of the system to prevent monitor evasion, as identified in Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure.
- Reasoning and Self-Evolution: DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement, GraphSkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback, and A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents provide the mathematical framework for credit assignment and skill distillation. To Think or Not to Think: Allocating Reasoning Where It Helps and CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment optimize the “compute tax” of deep reasoning.
Theme 3: Mechanistic Interpretability and Structural Analysis
To trust these systems, we must look under the hood. We are moving away from post-hoc explanations toward “mechanistic interpretability”—mapping the internal circuits that drive behavior.
- Circuit Discovery: Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models identifies specific attention-routing circuits responsible for hallucinations, allowing for targeted interventions.
- Structural Analysis: Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design reveals a “Global Positional Band” in Transformers, while Unraveling the cognitive patterns of Large Language Models through module communities maps LLM modules to cognitive skills, suggesting a distributed organization akin to biological brains.
Theme 4: Diagnostic-Driven Evaluation and Robustness
Aggregate metrics are insufficient for safety-critical systems. We are entering an era of granular, diagnostic-level auditing.
- Diagnostic Evaluation: Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling and Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation expose how aggregate accuracy masks failure modes. StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection provides statistical rigor for worst-case reliability.
- Robustness and Privacy: TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement and Beyond Average Safety: Chance-Constrained LLM Fine-tuning ensure safety-critical robustness. AIR: Analytic Imbalance Rectifier for Continual Learning and Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning address the challenges of learning in dynamic, privacy-sensitive environments.
Theme 5: Embodied Intelligence and Spatial Grounding
AI is stepping out of the screen and into the physical world, requiring a deep understanding of 3D space, causality, and temporal consistency.
- World Models: NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation and World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal allow robots to rehearse actions in “mental simulators.” Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners calls for a shift toward true causal reasoning.
- Spatial Reasoning: Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models, STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models, and Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning demonstrate that explicit geometric tools and object-centric monitoring are essential for reliable spatial reasoning.
- Efficiency and Interaction: Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs, Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use, and AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing optimize compute and state management for complex robotic and design tasks. PUBG Ally: A Conversational Embodied Agent as an AI Teammate and Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots explore the social dimensions of embodied AI.