ArXiV ML/AI/CV papers summary
Theme 1: World Modeling & Latent Dynamics
The field is shifting from static pattern recognition to the construction of “world models”—systems that learn the underlying causal, temporal, and structural dynamics of their environments. By capturing the “physics” of data, these models move beyond simple interpolation toward true simulation and reasoning.
- A Bayesian Mirror Architecture for Emergent Consciousness: Circular Hierarchies, Self-Manifolds, and Hybrid Event-Self Binding proposes a generative framework where sensory and self-latents interact through recursion, suggesting that consciousness is an architectural property of systems with circular structure.
- Directed Temporal Representations for Offline Visual Control introduces a geometry aligned with temporal reachability, allowing models to estimate the “cost” of reaching a goal rather than just predicting the next state.
- Towards Financial World Modeling establishes a foundation for financial world models by introducing a massive dataset (Market-1T) and a rigorous evaluation protocol to compare representation learning strategies.
- DSReg: Provably Recovering Individual World Latents without Reconstruction provides a theoretical breakthrough, proving that individual world latents can be recovered without decoders or labels, provided the system exhibits “Structural Diversity.”
- RoboJEPA: Scaling Robotic Latent World Models establishes scaling laws for robotic world models, showing that imagination error is a reliable proxy for real-world planning performance.
- A Review Of Robotic World Models For Dynamic Environments Based On Factor And Scene Graphs provides a comprehensive look at how these models are evolving to handle dynamic environments, emphasizing the need for hybrid models that combine geometric factor graphs with semantic scene graphs.
- Geometry-Centered 3D Latent World Models for Growing Surfaces emphasizes that for robots to operate in the real world, they must represent geometry explicitly rather than relying on pixel-space shortcuts.
Theme 2: Agentic Reasoning, Reliability, & Governance
As AI transitions from lab-based benchmarks to high-stakes deployment, the focus has moved toward “governance by design.” This involves ensuring agents remain within safe, verifiable bounds while managing the “scarcity inversion”—a state where reasoning is abundant, but trustworthy evidence and physical execution are the true bottlenecks.
- Bounded Autonomy and Verifiable Safety for Agentic AI Enabled Automation introduces BRaVeS, a framework that uses Lyapunov-bounded consensus to enforce safety constraints, reducing autonomy when epistemic risk is high.
- HydroSphere: A Framework for Governed, Self-Healing Wastewater Infrastructure applies these principles to critical infrastructure, combining reinforcement learning with explicit operational safeguards.
- GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents addresses the “drift” problem in agentic training by using dependency graphs to ensure that supervision is concentrated on the most impactful steps of a trajectory.
- Robust Decentralized Fairness Auditing tackles the challenge of auditing LLMs for fairness in a decentralized setting, protecting against “fairwashing” by adversarial auditors.
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning and RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement highlight how agents can distill raw experience into reusable skills, creating a virtuous cycle of improvement.
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams and Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents demonstrate that by analyzing the “trace” of an agent’s reasoning, we can provide exact, intermediate feedback.
- Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator and Questioning the Questions: Sustaining Self-Evolution in Reasoning Models show that agents can bootstrap their own reliability by training detectors against their own outputs.
- Routing-Aware Safety Alignment for Mixture-of-Experts Models and SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models explore how to align complex, sparse architectures without sacrificing general utility.
- SpecGuard: Proving a Task Is Broken Before the Agent Cheats introduces a rigorous approach to safety, using formal methods to certify that a task is not misspecified before an agent is allowed to attempt it.
- Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents highlights that agents are surprisingly poor at detecting silent failures, while The Handover Problem: Governing Autonomy Transitions in Human-AI Collaboration introduces the Handover Readiness Score (HRS) for escalating control back to humans.
- QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents and From Verification Failures to Reusable Guidance for Coding Agents demonstrate that agentic systems can be trained to perform formal verification.
Theme 3: Efficient Inference & Architectural Innovation
To support the next generation of AI, researchers are optimizing how models manage memory and compute. This includes smarter KV cache management, hardware-native architectures, and efficient long-context processing.
- KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression introduces a prefetching mechanism to recover verbatim copying in compressed caches.
- A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention uses adaptive decay to prune tokens, achieving up to 25x compression.
- SPIN: Shadow Predictive Indexer for Sparse Attention uses history-based prediction to identify important KV blocks, improving throughput.
- Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization demonstrates that compressing sequences into “block latents” improves generation throughput by over 6x.
- Attention via Black-Box Vector Search and Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference provide methods to scale attention to long contexts without quadratic costs.
- CNet: A Complex-Valued Deep Learning Framework with Wirtinger Autodifferentiation and FFT–Hadamard Convolution and An RRAM-based Hardware Implementation of a Radial Basis Function Neuron for Edge Classifiers explore non-traditional architectures that align better with physical hardware.
- TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning and Clean: Second-order LLM Training at Linear Memory Cost via Nystr"om Sketching demonstrate state-of-the-art training with a fraction of the memory overhead.
Theme 4: The Geometry of Learning & Representation
A recurring theoretical insight is that the “geometry” of a model’s internal representation is often more critical than the specific training objective. Understanding these geometric constraints is demystifying deep learning, moving it from “black magic” to a rigorous science.
- Are Parameter-Efficient Fine-tuning Methods Really Different? finds that PEFT methods work largely because they preserve pretrained weight geometry.
- Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data demonstrates that independently trained models often converge to a shared geometric structure.
- Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking shows that neural network “clusters” result from breaking hidden symmetries in the architecture.
Theme 5: Security, Robustness, & Adversarial Defense
As agents gain autonomy, they become targets for sophisticated attacks. Research is shifting toward “execution boundary” protection and robust defense mechanisms that survive adversarial manipulation.
- Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks and Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models highlight the ongoing cat-and-mouse game between providers and adversaries.
- Visual Hijacking: Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution and Visual Memory Attacks Can Persist Through The KV Cache show that adversarial inputs can persist in memory to steer agents.
- Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files demonstrates supply chain risks in agentic workflows.
- APEX: Active Protection at Execution Boundaries for LLM Agents proposes gating agent actions via pre-compiled authorization contracts.
- Certification of Real Images through Calibrated Content Authentication and Protective Perturbations Must Survive the Resize: Scale-Robust Image Immunization against Malicious Editing tackle deepfake detection and image protection.
Theme 6: Embodied AI, Robotics, & Vision-Language-Action (VLA)
The frontier of AI is shifting toward physical interaction, where language, vision, and action are tightly coupled. These models must ground their reasoning in 3D space to operate effectively in the real world.
- SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation and S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens enable real-time 3D scene exploration and persistent state integration.
- Agentic Scene Policies and Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies use iterative feedback to bridge the gap between simulation and reality.
- Video Prediction Policy 2: Predict Better, Act Better and SpanVLA: Learning from Negative-Recovery Samples with Fast Action Bridging for Vision-Language-Action Model explore long-horizon planning and failure recovery.
- Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving forces models to ground reasoning in 3D space before planning.
- Juno: Taming Predictive Latents for Vision-Language-Action Models aligns VLA models with embodiment-specific control.
- PhysEvo: Astra Can Act, Let It and EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution demonstrate autonomous diagnosis of physical failures.
- OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning emphasizes explicit geometric representation.
Theme 7: Scientific Discovery & Domain-Specific AI
AI is evolving into a specialized “AI Scientist,” capable of discovering causal drivers and building models in complex fields like climate science, biology, and medicine.
- SciExam for ENSO: Can AI Agents Build Climate Models? and DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists show agents outperforming human baselines in scientific discovery.
- Fourier Feature Pyramids for Physics-Informed Neural Networks and Fluids You Can Trust: Property-Preserving Operator Learning for Incompressible Flows embed physical laws directly into architectures.
- CircuitATLAS: Agentic reasoning over a systems neuroscience knowledge graph for target discovery in circuitopathies and When Scientific Cognition Is No Longer Scarce explore the “scarcity inversion” in scientific research.
- AI-Assisted Computational Reproducibility on the FABRIC Testbed and Comprehension Audits to Mitigate Risks from Automated AI Research address the need for human oversight in AI-driven scientific progress.
- Zero-Shot Brain MRI Inpainting with 2.5D Unconditional Flow Priors and Unified Multi-plane Autoregressive Diffusion for 3D Multi-contrast MRI Synthesis improve medical diagnostic reliability through generative synthesis.
- One Frame, Full Heartbeat: ECG-Free 4D Cardiac Cine MRI Synthesis via Radial-Decomposed Flow Matching and Heartian: Physiology-Aware Relightable Gaussian Head Avatar incorporate physiological priors into simulations.
- Quantifying Volumetric Risk: Class-Aware Asymmetric Weighted Conformal Prediction for 3D Medical Image Segmentation provides statistical guarantees for clinical AI.
Theme 8: Generative Models & 3D Gaussian Splatting
3D Gaussian Splatting (3DGS) is being optimized for efficiency and control, with new tools emerging to automate the translation of research into actionable code.
- SPLATIFY: Reproduce, Discover, Innovate! From Papers and Ideas to Trainable 3DGS Code uses multi-agent frameworks to accelerate 3DGS innovation.
- RDGSplat: Render-Dedicated Geometry for Novel View Synthesis and DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting optimize geometry and pose estimation.
- TileSkipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting and Efficient 3D Gaussian Head Avatars for Edge Devices enable high-quality rendering on resource-constrained hardware.