ArXiV ML/AI/CV papers summary
This collection of research highlights a pivotal shift in machine learning: we are moving away from “brute-force” scaling toward a more nuanced, structural understanding of how models learn, store, and retrieve information. The field is maturing from an era of “black-box” scaling into an era of agentic precision, where systems are modular, verifiable, and capable of self-correction.
Theme 1: Structural Efficiency & Sparse Computation
The quest to make large models deployable on resource-constrained hardware is no longer just about shrinking parameters; it is about rethinking the fundamental operations of inference.
- Sparsity & Gating: OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit and Resource-Aware Federated Mixture-of-Experts with Adaptive Pruning for Onboard Learning in LEO Satellite Constellations optimize expert utilization. Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs and GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning demonstrate significant FLOP reductions through smarter sparsity.
- Quantization: Product-Aware Deterministic Rounding for Quantized Matrix Multiplication, Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights, and Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter push quantization to extreme limits without catastrophic accuracy loss.
- Inference Optimization: Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders, FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward, and Spexis: Speculative Lookahead Scheduling for LLM Inference rethink data access and parallelism to accelerate generation.
Theme 2: Mechanistic Interpretability & Representation Geometry
We are increasingly treating neural networks as “experimental systems,” mapping the internal geometry of models to understand how they store facts and make decisions.
- Internal Dynamics: How Linear Attention Remembers and Counting on Thinking: Tracing Evidence Integration in Language Models explore how models allocate “thinking” time. Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction reveals hidden functional structures often dismissed as noise.
- Geometry & Logic: The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning and Intuition vectors demonstrate that reasoning is a geometric property of the representation space.
- Steering & Control: Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs and Activation Flow: Manufacturing Activations for Steering allow us to intervene directly on internal activations to guide model behavior.
Theme 3: Agentic Reasoning & Self-Evolution
The field is transitioning from passive text generation to autonomous agents that can plan, verify, and improve their own capabilities.
- Self-Evolution: Frontier Learning: Training LLM Reasoners at the Edge of Capability and RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation highlight a shift toward open-ended, self-improving training. Program-Verified Self-Evolution for Vision-Language Models uses structured programs to verify facts during self-training.
- Agentic Frameworks: LLMs are General Asynchronous Agents and ActionEngine: From Reactive to Programmatic Web Agents via State Machine Memory address the practical realities of agentic workflows, including asynchronous inputs and stateful memory.
- Verification: PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety and Certified Multi-Source Integrity for Structured Agent Actions provide frameworks to ensure that high-stakes actions are based on verified evidence.
Theme 4: Physics-Informed & Scientific Machine Learning
Machine learning is increasingly being used to solve complex physical systems by embedding fundamental laws—geometry, topology, and physics—directly into the architecture.
- Physics-Informed Surrogates: Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It and Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency enforce physical constraints during training.
- Scientific Discovery: D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders and MultiEcho: An Experimental Science of Learned Worlds treat world models as experimental systems for counterfactual interventions.
- Geometric Symmetries: Discovering Symmetries in Neural Network Parameter Spaces and On Parameter Symmetries and Conservation Laws in Gradient Flow provide a unified geometric framework connecting network symmetries to physical conservation laws.
Theme 5: Evaluation, Reliability & Alignment
As models become more capable, we are moving away from simple accuracy metrics toward “closed-loop” and “adversarial” evaluations that better reflect real-world deployment.
- Robustness & Auditing: ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models and LLM Unlearning Evaluation with TRIAGE provide frameworks to audit models for hidden vulnerabilities and ensure that “unlearning” actually removes sensitive information.
- Statistical Foundations: A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions frames distillation as a Bayesian process, moving the field from “engineering recipes” to a rigorous science of model alignment.
- Closed-Loop Evaluation: What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators argues that standard metrics often mask critical failure modes in real-world deployment.