ArXiV ML/AI/CV papers summary
Theme 1: Physics-Informed & Geometry-Aware Modeling
We are witnessing a profound shift away from purely statistical correlation toward models that respect the fundamental laws of our universe. By embedding physical and geometric constraints directly into the architecture, we move from “black-box” approximation to reliable, interpretable scientific discovery.
- Physics-Informed Neural Networks (PINNs): Physics-informed neural networks by Gradient-Guided Gaussian Adaptive Sampling (3GAS-PINNs) utilizes adaptive sampling to resolve intermittent structures like shock waves, yielding a 14x accuracy boost.
- Geometric & Topological Constraints: Spectral origin of the topological gap exponent d + {\eta}: mechanism, kernel, decomposition, and scope provides an analytical foundation for topological persistence, while DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity optimizes hardware-native inference by jointly learning logic functions and connectivity.
- Scientific Surrogates & Grounding: Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate and Development and Validation of a Physics-Guided Machine Learning Extrapolation Framework Using a Classical Transient Diffusion Benchmark demonstrate that physics-guided architectures enable robust extrapolation. Similarly, Self-Evolving Scientific Agent Designs Physically Reasoned White-Box Fluid Control and PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal T2 mapping show that injecting domain-specific laws—like fluid dynamics or Bloch equations—is essential for high-fidelity scientific modeling.
- Spatial Intelligence: GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion uses a “Geometry-Then-Appearance” paradigm to prevent photometric collapse. This geometric focus extends to StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning, EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning, and KODAMA: Multimodal Digital Twin Reconstruction for Urban RF Propagation Modelling, all of which prove that understanding 3D structure is the key to true spatial intelligence. Conversely, Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception highlights the persistent gap between linguistic fluency and physical intuition.
Theme 2: Agentic Reasoning, Tool Use, and Verification
The field is transitioning from passive chatbots to autonomous agents capable of orchestration, planning, and self-correction. As these agents take on high-stakes roles, the focus has shifted from mere generation to verifiable reasoning and artifact integrity.
- Agentic Frameworks: We are moving toward “spec-first” development. Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches and A-JIT: Agentic Just-In-Time Software Construction emphasize engineering discipline, while JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition enables coordination across heterogeneous environments. World-Time Compute with Verified Code World Models and Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization further show how agents learn search strategies through executable practice.
- Reliability & Verification: RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement, Verify to Amplify: Improving Reasoning via Learned Chain-of-Thought Verification, and CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots prioritize the verification of reasoning paths. Quantifying Logical Consistency in Transformers via Query-Key Alignment allows us to look inside the “black box” to detect when reasoning diverges from decision-making.
- Artifact Integrity: Can AI Agents Detect and Repair Artifact Drift in Network Experiments? and EvidenceNet: A Runtime Assurance Layer for Deciding whether Coordinated Agent Operations have Achieved an Operator’s Network Intent introduce “artifact integrity,” ensuring that agent claims remain supported by evidence throughout a task’s lifecycle.
- Agentic Benchmarking: OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows and GauntletBench: Re-evaluating the Capabilities of Agents Beyond Familiar Environments argue that we must evaluate the process of reasoning, not just the final output.
Theme 3: Memory Lifecycle and Multimodal Foundation Models
Persistent agents must manage their “mental” space like a living library, while foundation models are evolving toward a more native, holistic understanding of multimodal data.
- Memory Management: What Should an Agent Forget? Separating What Is Stored from What Is Used and Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs address the “boundless storage” problem through behavioral ontologies. ConvMem: Convolutional Memory for Long-Context Reasoning further optimizes this by reformulating long-context reasoning as a hierarchical convolution.
- Multimodal Evolution: Let ViT Speak: Generative Language-Image Pre-training simplifies vision-language alignment by predicting language tokens directly from visual tokens. To handle long-form tasks, VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning, Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs, and Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation propose hierarchical and tool-based approaches to maintain consistency over time.
- Embodied AI: LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation and Show-Harness: Just a VLM Agent Can Play Robots bridge the gap between high-level intent and low-level physical action.
Theme 4: Efficiency, Trust, and Accountability
As AI integrates into critical infrastructure, we must balance performance with privacy, fairness, and rigorous mathematical optimization.
- Efficient Inference: Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution, EFQ-Softmax: Exp-Free Quantization for Softmax, and Forward-Free LLM Depth Pruning via Weight Redundancy achieve near-lossless performance with reduced memory footprints. PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling optimizes on-device power consumption.
- Privacy & Fairness: Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning and Privacy-Preserving Split Learning for Federated LLM Fine-Tuning address privacy leakage. Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction and XAI-Refine: An Automated Explanation-Knowledge Loop for Brain-Age Prediction provide structured auditing tools.
- Robustness & Uncertainty: Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks and Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending emphasize multi-attack resilience. Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes, Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation, Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models, and CGSM: Concept-Guided Segmentation Model for Precise Pulmonary Lesion Delineation ensure reliability in high-stakes diagnostics. Towards Characterizing Scientific Image Utility and Upgradability further establishes a framework for detecting scientific inaccuracies.
- Optimization Theory: Variance-Reduced Fast Krasnoselkii-Mann Methods for Finite-Sum Root-Finding Problems, New Accelerated Past-Extragradient Methods with Variance Reduction for Generalized Equations, RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers, and TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation provide the mathematical and architectural infrastructure for faster, more stable training.