ArXiV ML/AI/CV papers summary
Theme 1: Physics-Informed and Geometry-Aware Modeling
The field is undergoing a fundamental shift from purely statistical “black-box” models to architectures that respect the underlying laws of the universe. By embedding physical constraints and geometric priors directly into neural networks, we ensure that AI outputs remain consistent with reality—whether that means preventing “uphill water” in flood models or maintaining structural integrity in engineering simulations.
- Physics-Guided Surrogates: SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics, COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations, and Physics-Informed CNN-LSTM for Street-Scale Urban Flood Prediction: Reconciling Aggregate Accuracy and Street-Level Plausibility integrate governing equations into training objectives to ensure physical consistency. Similarly, Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision and FreeLit: Paired-Free Indoor Relighting via Physics-Guided Diffusion use physical priors to guide generative processes.
- Geometric and Scientific Constraints: Generative Distributionally Robust Optimization and Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning leverage data geometry for optimization. In the sciences, A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes, OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation, and Group Equivariant Diffusion for Anomaly Detection in Computational Cytology demonstrate how domain-specific knowledge and symmetry constraints lead to more reliable, scientifically valid outputs.
- 3D Reconstruction: InnerGS: Internal Scenes Reconstruction and Segmentation via Factorized 3D Gaussian Splatting and Diff2DGS: Reliable Reconstruction of Occluded Surgical Scenes via 2D Gaussian Splatting move beyond surface-level modeling to capture internal structures and complex deformations.
Theme 2: Agentic Systems and Workflow Compilation
We are moving beyond simple “chatbots” toward autonomous agents that act as reasoning partners. This transition involves treating natural language not just as context, but as executable code that requires orchestration, verification, and long-horizon planning.
- Workflow Compilation & Planning: COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution and ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design treat instructions as source code to be compiled into control-flow graphs. Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering, HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising, ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning, and OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation focus on hierarchical decomposition and the orchestration of multi-agent systems.
- Tool-Augmented Reasoning: SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task, Addressable Recall Compaction for Long Context-Window Control in AI Agents, ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition, Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing, and AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities demonstrate how agents decompose complex tasks by selecting and composing specialized tools.
Theme 3: Trust, Verification, and Epistemic Governance
As agents gain autonomy, we must shift from blind trust to verifiable claims. This theme emphasizes building systems that are auditable by design, ensuring that model rationales are secondary to objective, server-verified evidence.
- Verification Frameworks: Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales provides a mediation layer for verification. Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines and Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction argue that verification is a necessary component of responsible human-AI collaboration.
- Uncertainty and Selective Prediction: FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting, Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference, Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks, and Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification provide methods to quantify uncertainty and abstain from predictions when confidence is low.
Theme 4: Efficiency, Adaptation, and Edge Deployment
To make AI truly ubiquitous, we must overcome the “scale-at-all-costs” bottleneck. This research focuses on compressing models, optimizing inference, and enabling efficient adaptation to new tasks without the prohibitive cost of full-stack retraining.
- Quantization and Sparsity: Stable FP4 Training via Transposition-Invariant Block Quantization, GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization, Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising, DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers, and Mondrian: On-Device High-Performance Video Analytics with Compressive Packed Inference push the limits of low-precision and efficient inference.
- Efficient Architectures and Adaptation: FFNet: MetaMixer-based Efficient Convolutional Mixer Design, CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model, REPREC: Representation Driven Parameter-Efficient Recommendation System, OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs, SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models, OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models, and Localized Adaptation Reveals Distinct Learning Signatures in Transformers demonstrate how to achieve high performance through targeted updates and token compression.
Theme 5: Interpretability, Security, and Statistical Rigor
The final theme addresses the “black box” problem through mechanistic interpretability and robust security auditing, while grounding the field in rigorous statistical foundations.
- Interpretability and Auditing: Interpretable GOHR Agents via Sparse Autoencoders, Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation, Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation, Detecting CSAM Text-to-Image LoRAs From Weights, and Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering provide tools to audit model internals and prevent harmful content.
- Security and Robustness: Distributing Security Controls Through Harness Engineering, Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks, Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection, and Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study highlight the need for external governance and advanced detection techniques.
- Statistical Foundations: polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs, Generalised Robust Bayes for Joint Inference of Model and Contamination, and Conformal changepoint localization reinforce the mathematical rigor required for reliable machine learning.
- Clinical Intelligence: Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation, Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management, PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents, MEDIC-AD: Towards Medical Vision-Language Model’s Clinical Intelligence, PRIMA: Pre-Training with Risk-Integrated Image–Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation, and Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs demonstrate the application of these principles to high-stakes medical and safety-critical environments.