ArXiV ML/AI/CV papers summary
This collection of research represents a pivotal moment in machine learning, where the focus is shifting from simply scaling models to understanding the mechanics of intelligence—how models store information, how they reason, and how they interact with the physical world. Like the great astronomers who mapped the heavens to understand the laws of gravity, these researchers are mapping the “latent space” of neural networks to understand the laws of computation.
Theme 1: The Mechanics of Reasoning and Memory
We are moving away from “black-box” models toward systems that can be inspected, audited, and controlled. Reasoning is increasingly viewed as a deliberate, costly allocation of computation rather than an inherent property of model weights.
- Counting on Thinking: Tracing Evidence Integration in Language Models shows that LLMs use “thinking” tokens to perform tasks they cannot handle automatically.
- The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning introduces STRAT, an architecture that partitions the residual stream into “Data” and “Type” subspaces to mimic human logical separation.
- How Linear Attention Remembers provides a mechanistic look at how linear attention models store information, highlighting the capacity limits imposed by cross-fact interference.
- Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations and Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models challenge the idea that activations alone explain model behavior, proposing “knowledge matrices” as more fundamental units of analysis.
- Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths and Causal Routing for Unlearning demonstrate that true “forgetting” requires severing the relational reasoning paths that reconstruct facts.
Theme 2: Efficient Scaling and “Frugal” Intelligence
As models grow, the “memory wall” and inference costs have become planetary-scale challenges. The community is shifting toward “smarter” rather than “larger” systems.
- OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit and SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving use signal reconstruction and dynamic masking to prune redundant experts.
- NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization proves that engineering the training pipeline can allow a 6.5M-parameter model to outperform 200M-parameter giants.
- SketchSSM: Write to the Full State, Read from a Compact Sketch and FlashSampling: Fast and Memory-Efficient Exact Sampling reduce memory access traffic and overhead in state-space models.
- ElasticKV: A Mixed-Fidelity KV Runtime for LLM Serving and PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding introduce runtime-managed fidelity to handle massive context lengths.
- Chameleon: Dynamic Format Adapter for Efficient Diffusion and EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates push the limits of quantization by treating the format itself as a variable.
Theme 3: Physics-Informed and Scientific AI
AI is increasingly being applied to the physical sciences, where models learn the underlying differential equations that govern our universe rather than just predicting patterns.
- Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It and Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency ensure models respect governing physical equations.
- Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning and Dynamic Kuramoto-Hodge Operators for PDEs on Complex Geometries and Topologies encode topological constraints to achieve superior accuracy with fewer parameters.
- MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception and Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer inject geometric knowledge to solve complex physical tasks with less data.
- The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning offers a sobering critique, reminding us that we must encode physical structure, not just minimize loss.
- MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws? and Autonomous phase discovery demonstrate that AI can autonomously discover physical laws and quantum phases.
Theme 4: Agentic Reasoning and Self-Evolution
We are moving from models that predict the next token to agents that plan for the next outcome, manage memory, and align with human intent through verifiable rewards.
- RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation and RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation show that agents improve significantly when forced to verify outputs against formal constraints.
- Entropic Advantage Policy Optimization (EAPO) and GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents tackle the “sparse reward” problem in long-horizon tasks.
- OPT-Zero: Learning to Optimize through Solver-Grounded Self-Play and GenMem: Generative Symbolic Memory for Self-Evolving Harness represent a paradigm shift: agents that generate their own tasks and manage their own symbolic memory.
- Program-Verified Self-Evolution for Vision-Language Models replaces noisy consensus with structured programs to ensure self-evolving models train on ground-truth facts.
- A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security establishes a foundational reference architecture for the agentic ecosystem.
Theme 5: The Ethics of Evaluation and Robustness
The community is developing a more rigorous “scientific method” for AI evaluation, moving past simple accuracy metrics toward understanding why models succeed or fail.
- ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models and The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection highlight that current benchmarks are often fragile and require standardized, pre-registered audits.
- When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning reveals that reinforcement learning can make models wildly inconsistent in reporting their own failures.
- ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control and PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety propose runtime guardrails that intervene in real-time to prevent irreversible harm.
- One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs and COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails address the persistent vulnerability of multimodal models to adversarial perturbations.