ArXiV ML/AI/CV papers summary
Theme 1: Physics-Informed and Embodied Intelligence
We are witnessing a transition from abstract, data-hungry models to systems that respect the fundamental laws of the universe. By embedding physical constraints and geometric awareness into neural architectures, we move closer to machines that truly “understand” the world they inhabit.
- Physics-Informed Learning: Innovations like Lecture notes on Physics Informed Neural Networks, Neural Operators, and their applications, HiLNO: A Hierarchical Latent Neural Operator with Multi-Scale Supervision for PDEs on General Geometries, and Physics-Informed Neural Networks for Fast Multilayer Spectral Inversion of H{\alpha} 6562.8 A and Ca II 8542.1 A Spectra demonstrate how incorporating physical laws reduces computational costs for complex simulations. Robustness is maintained through techniques like Stable Filters for Generative Modeling of Graph Signals and Tackling Failure Modes of PINNs and PIKANs Using Conflict-Free Gradients.
- Embodied AI: Moving beyond screens, AI is becoming “situated.” Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance and Visual Cue Guided Video Planning for Generalizable Robot Navigation ground language in physical action. To bridge the “sim-to-real” gap, researchers are utilizing WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories and Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents.
- Physics-Grounded Generation: New models are learning to simulate reality, as seen in PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control and LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting.
Theme 2: Agentic Reasoning and Reliability
As AI evolves from passive chatbots to autonomous agents capable of long-horizon planning, the focus has shifted toward “agentic reliability”—ensuring these systems can reason, abstain when uncertain, and operate within safe, contract-based boundaries.
- Failure Detection and Calibration: Agents must know their own limits. Locating Hidden Failures Makes Long-Horizon Agents More Reliable and The Missing “I Don’t Know”: Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention emphasize the necessity of calibrated uncertainty.
- Structured Reasoning: To prevent “hallucinations,” researchers are decomposing tasks into interpretable stages, as explored in A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning and Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents.
- Governance and Safety: Autonomous systems require oversight. Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement and Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery highlight the risks of agentic collaboration. Solutions include Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records, Symbolic Temporal Supervision of LLM Agents Using Contracts, and Visual Compliance via Executable Safety Rule Entailment.
Theme 3: Efficient Inference and Optimization
The “memory wall” is the great bottleneck of our time. To deploy trillion-parameter models, we must move toward smarter, more efficient architectures that prioritize hardware-aware acceleration and adaptive computation.
- KV Cache and Memory Management: Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches and GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference optimize long-context performance.
- Quantization and Pruning: We can achieve efficiency through Higher-order pruning of experts in mixture-of-experts language models and Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing. Hardware-specific gains are found in FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs and The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction.
- Adaptive Computation: Not every query requires the same compute. One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG, Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition, StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection, and VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs demonstrate dynamic resource allocation.
Theme 4: Multimodal Reasoning and Interpretability
Modern AI must bridge the gap between visual perception and logical reasoning. This requires “opening the black box” to ensure that models are not just pattern-matching, but truly understanding the data.
- Spatial and Multi-Hop Reasoning: PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection and RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection improve spatial precision. Complex reasoning is addressed in MechReason: Benchmarking Multi-Image Multi-Hop Reasoning in Mechanical Engineering, Multi-Hop Knowledge Composition is Bound by Pretraining Exposure, and Reasoning with Image Generation.
- Interpretability and Causal Auditing: We must understand why models act. NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections and Regional Explanations via Causal Sufficiency and Necessity provide functional insights, while No Usable Linear “Capitulation Direction” in Two Small LLMs and Steering Interference Reflects the Model’s Defaults, Not the Behavior Directions offer cautionary perspectives on model steering.
- Hallucination and Security: SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation, Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models, and Unsafe by Reciprocity: How Generation-Understanding Coupling Undermines Safety in Unified Multimodal Models address the critical intersection of safety and multimodal integration.
Theme 5: Scientific Discovery, Evolution, and Society
AI is becoming a partner in scientific inquiry, while simultaneously forcing us to confront the biological and societal implications of our creations.
- Scientific Agents: AI is accelerating discovery in Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents, ScienceIDE: Turning World’s Scientific Codebase into Agent Learnable Environments, OphthaReason, and SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis. Medical imaging advances include A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data, NeuroTS-Net: Multi-Class Semantic Segmentation of Pediatric Brain Tumors in Multi-Modal MRI, and MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis.
- Evolutionary Dynamics: We are beginning to view AI through the lens of population genetics, as seen in The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations, Agora: Git as Shared Memory for Collective AutoResearch, and PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research.
- Societal Impact and Evaluation: We must address bias in From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale, Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment, Which Demographics do LLMs Default to During Annotation?, DenseFace: Bias Mitigation in Face Recognition via Density-Aware Probabilistic Matching, and ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models. Ultimately, we must prioritize Human Resilience in the AI Era – What Machines Can’t Replace and adopt rigorous benchmarks like TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue, EgoPathBench: Evaluating Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models, MiRAGE: Evaluating Multimodal Retrieval Augmented Generation, VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal, and Code Consistency Preference Optimization Verification for Language Model Alignment.