ArXiV ML/AI/CV papers summary
Theme 1: Mechanistic Interpretability and Geometric Foundations
The “black-box” era of deep learning is yielding to a new paradigm of structural transparency. By treating neural networks as geometric objects, researchers are uncovering the “laws” of representation that govern how models store concepts and make decisions.
- Geometric Tracking: Capsule Lens: Locating and Tracking Concept Geometry in Model Representations provides a rigorous framework for observing how post-training restructures a model’s internal space.
- Structural Interpretability: PhysSAE: Mechanistic Interpretability with Sparse Autoencoders and “World Knowledge” in the Weights: Reading Concept Circuits of Vision Transformers demonstrate that we can causally interrogate latent representations to identify specific “concept circuits.”
- Diagnostic Insights: Medical AI Encodes a “Feeling of Error”: Verifying Cancer Segmentation via Internal Concepts highlights how internal model states can signal uncertainty before an error manifests, while GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks provides tools to systematically localize failure modes.
- Linear Algebra Foundations: Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression offers a theoretical basis for why specific compression techniques succeed where others fail by analyzing the natural rank collapse of transformers.
Theme 2: Efficiency, Sparsity, and Adaptive Computation
We are moving toward “frugal AI,” where performance is maximized through architectural intelligence rather than brute-force parameter scaling.
- Extreme Compression: Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity and All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs push the boundaries of model size, while Squeeze10-LLM: Squeezing LLMs’ Weights by 10 Times via a Staged Mixed-Precision Quantization Method and SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions provide frameworks for simultaneous pruning and quantization.
- Adaptive Routing: ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs, Hyperparameter Scaling Laws Across MoE Sparsity, and EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs demonstrate how to virtualize and consolidate experts to overcome memory bottlenecks.
- Reasoning-Aware Efficiency: Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment warns that uniform quantization can harm reasoning, advocating for the protection of critical circuits.
Theme 3: Agentic Reasoning, Alignment, and Governance
The frontier has shifted from static text generation to autonomous, goal-directed agents. This transition necessitates new frameworks for reliability, tool-use, and recursive self-improvement.
- Reasoning & Planning: A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM and Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR formalize “slow thinking” in AI. ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR optimizes the rollout process for verifiable rewards.
- Self-Evolution: MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves, AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents, and NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness showcase systems that autonomously refine their own training machinery.
- Governance & Safety: PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations and A Unified Policy Architecture (UPA): The Governance Kernel for Enterprise AI Operating Systems propose formal frameworks for managing authority and auditability. The Normalization of Deviance in AI Development serves as a sobering reminder that safety is as much an organizational challenge as a technical one.
Theme 4: Physics-Informed and Embodied Intelligence
AI is increasingly acting as a “scientific instrument,” embedding physical laws into neural architectures to solve complex problems in fluid dynamics, robotics, and biology.
- Scientific Operators: Local gradient neural operator, Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics, and SFO: Learning PDE Operators via Spectral Filtering demonstrate how neural operators respect underlying physical laws.
- Embodied World Models: PAN: A World Model for General, Actionable, and Long-Horizon World Simulation and GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation highlight the shift toward models that understand physical causality and counterfactual outcomes.
- Geometric Grounding: LightSplat: Real-Time High-Fidelity 3D Gaussian Splatting with Loop Closure and TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid bridge the gap between 2D pixels and 3D physical reality, ensuring that AI-driven discoveries and simulations are geometrically valid.
Theme 5: Federated Learning and Robustness
As data privacy and security become paramount, the field is developing decentralized protocols that ensure models remain auditable and resilient to poisoning.
- Decentralized Learning: Revisiting One-Shot Federated Graph Learning: Training-Free Statistical Estimation and Trust-But-Verify: Poisoning-Resilient Locally Private Graph Learning Protocols provide robust methods for training across heterogeneous devices without centralizing sensitive information.
- Auditable Deletion: Can an AI Assistant Really Forget? Auditable Deletion from Addressable Memory explores the technical challenges of “machine unlearning,” proposing memory gates that allow for the verifiable removal of specific records.