ArXiV ML/AI/CV papers summary
As we stand at the intersection of machine learning, physics, and decision science, we are witnessing a profound shift: we are moving away from “black-box” optimization toward systems that are structurally aware, physically grounded, and auditable. Much like the celestial mechanics that govern the orbits of planets, the algorithms presented here are increasingly governed by the “laws” of their domains—whether those laws are the conservation of energy in a fluid, the causal structure of a clinical trial, or the logical constraints of a symbolic reasoning task.
Theme 1: Physics-Informed and Geometry-Aware Learning
The most striking development is the move toward architectures that treat the laws of nature not as suggestions, but as hard constraints. We are no longer just training models to fit data; we are training them to respect the universe.
- Physical Constraints: Physics-informed reduced-order modelling with equivariant spectral submanifolds and TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning demonstrate that incorporating symmetry and physical structure provides robustness that pure data-fitting cannot. Guarantees by Construction for Learned Finite Volume Schemes on Steady Supersonic Flow moves from “soft” penalties to “hard” constraints, ensuring models cannot output physically impossible states.
- Geometric Embeddings: By embedding models in the correct geometric spaces—such as Lie groups or tropical manifolds—we achieve higher fidelity with less data, as seen in Tropical Algebraic Geometry for Neuronal Representations: An Arakelov-Green Measure Based Descriptor for Graph Learning and SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups.
- Neural Rendering & Physics: ACA-GS: Adaptive-Capacity Anchored Gaussian Splatting for Compact Dynamic Radiance Fields and Generative neural physics enables quantitative volumetric ultrasound of tissue mechanics show how coupling generative networks with physics-informed simulation enables high-fidelity reconstruction and analysis.
Theme 2: The Evolution of Agentic Reasoning and Persistence
We are moving from the “Big Bang” era of scaling parameters to a “Stellar Evolution” era, where we understand the internal life-cycle of models. LLMs are evolving from passive text generators into persistent, autonomous agents.
- Agentic Harnesses: Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) and OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents highlight the need for structured environments to decompose long-horizon tasks. DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models and Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens bridge the gap between abstract reasoning and concrete perception.
- Memory and Persistence: A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS), CogniFold: Always-On Proactive Memory via Cognitive Folding, and ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory treat memory as an evolving cognitive structure rather than a static lookup table.
- Strategic Planning: Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection warns that agents must look beyond immediate information gain to build long-term abstractions.
Theme 3: Reliability, Auditing, and Verification
As models enter high-stakes environments, the “vibe-test” is being replaced by rigorous, auditable frameworks. Verification is no longer an afterthought; it is a structural requirement.
- Auditing and Trust: Manipulation-Proof Oblivious Audits against Deceptive Model Providers and Local Violation Certification for Linear Predict-Then-Optimize Pipelines provide mathematical tools to audit models. From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs and Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems formalize intuitive “vibes” into decision-theoretic metrics.
- Structural Verification: The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents and SafeCommit: Certifying When Memory-Grounded Agents May Safely Act ensure agents act only under validated constraints. EviGraph: Evidence-Guided Autonomous Research Agents and Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning ground conclusions in explicit evidence graphs.
- Psychometric Evaluation: Item Response Theory for AI Safety and Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation provide the tools to measure model behavior under stress and quantify the degradation of validity across pipelines.
Theme 4: Embodied Intelligence and Domain-Specific Grounding
General-purpose models are increasingly being adapted for high-stakes, real-world interaction, requiring specialized grounding in robotics, medicine, and finance.
- Embodied AI: Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching, MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight, and GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking demonstrate the shift toward active, multi-modal interaction in dynamic environments.
- Clinical and Professional Grounding: MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays, FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables, and RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care emphasize that professional-grade AI must meet the tacit standards of human practitioners.
Theme 5: Efficiency and the “Physics” of Training
Efficiency is not just about smaller models, but about smarter training dynamics and mechanistic understanding.
- Training Dynamics: The Hamilton-Jacobi Theory of Deep Learning unifies training with PDE theory, while Neural Diversity Regularizes Hallucinations in Language Models and Regularization can make diffusion models more efficient show that diversity and sparsity improve performance.
- Speculative Inference: SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts and Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference use lightweight approximations to guide heavy model lifting.
- Mechanistic Insights: IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers and Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability provide a profound look at the fundamental limits of learnability and the internal mechanics of neural networks.