ArXiV ML/AI/CV papers summary
Theme 1: Efficient Inference and Architectural Innovation
The “memory wall” and the quadratic cost of attention are the primary bottlenecks in scaling AI. We are moving away from brute-force scaling toward surgical efficiency, ensuring that models can run on edge devices without sacrificing performance.
- Quantization and Pruning: PRQuant: Permutation Residual Quantization for Low-Overhead Inference and Correlation-Aware Structured Pruning for Large Language Models optimize weight storage and processing, while SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models and GPLQ: A General, Practical, and Lightning QAT Method for Vision Transformers push models into extreme low-bit widths.
- KV Cache and Token Optimization: To handle long contexts, PAGE: Partition-Aware Gated KV-Cache Eviction and StepKV: Step-Aware KV Cache Compression for LLM Agents manage memory bottlenecks. Similarly, VPRune: Efficient Training-free Pre-LLM Visual Token Pruning and PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference ensure we only process the most relevant visual data.
- Architectural Shifts: Scalable Mamba-Based Message-Passing Neural Decoder for Error-Correcting Codes and RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models explore alternatives to dense attention, while TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers and Accelerating Dense LLMs via L0-regularized Mixture-of-Experts surgically restructure models for speed.
Theme 2: Mechanistic Interpretability and Causal Grounding
We are entering the “neuroscience of AI,” moving beyond black-box evaluation to map the internal “geography” of models. This allows us to distinguish between post-hoc rationalization and true causal reasoning.
- Tracing Circuits and Concepts: Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs and From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models map how models store and manipulate knowledge. Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities and Toward a Unified Mathematics of Concepts suggest a universal geometry of semantics.
- Steering and Faithfulness: Read-Best Is Not Steer-Best: A Probing–Steering Layer Dissociation in Omni-Modal Large Language Models and Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models refine how we control model behavior. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness and How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation rigorously test if a model’s reasoning is actually driving its output.
- Quantum and Xeno-Interpretability: Watching Quantum Models Think: Hilbert-Space Interpretability in Quantum Transformer Blocks and Xeno-Interpretability: Investigating the Alien Minds of LLMs explore non-human internal representations.
Theme 3: Agentic Autonomy and Reliable Reasoning
The field is shifting from static text generation to autonomous agents that plan, use tools, and self-correct. The challenge is ensuring these agents are reliable, verifiable, and cost-effective.
- Agentic Workflows: AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows and Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving optimize resource usage. Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents and Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents focus on long-horizon memory and state management.
- Self-Refinement and Verification: DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement and Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning enable agents to learn from their own history. LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers and SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? use formal logic to ensure reasoning is sound.
- Safety and Reliability: Beyond Task Completion: Training Capable and Safe Computer-Use Agents, TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling, and XYEval: Agents say yes to bad advice address the critical need for guardrails in agentic systems.
Theme 4: Embodied AI and Physical World Models
We are moving toward “World Action Models” (WAMs) that integrate vision, language, and physical action. By grounding AI in geometry and tactile feedback, we enable robots to interact with the world with precision.
- Geometry and Physics: GAPS: Generative Active Pseudo-view Selection for Sparse-View 3D Gaussian Splatting and Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene anchor generation in 3D space. PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections and Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation enforce physical laws for real-world stability.
- Tactile and Multimodal Control: N0-Foundation: Towards the Age of Tactile Intelligence, Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control, and DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation integrate high-bandwidth tactile sensing. HuRo: Robotizing Human Videos for Scalable VLA Pretraining and Scaling Sim-to-Real VLA Reinforcement Learning with Generative 3D Worlds bridge the gap between simulation and reality.
Theme 5: Scientific Machine Learning and Domain-Specific AI
AI is becoming a “scientist-in-the-loop,” accelerating discovery in physics, medicine, and engineering by embedding domain constraints directly into neural architectures.
- Physics-Informed Discovery: StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting, Adaptive Physics-Informed Neural Networks for the Blasius Boundary-Layer Problem, and SDC-GON: Singular Decomposition and Consistency-Regularized Green’s Operator Networks for Solving Partial Differential Equations solve complex equations faster than traditional numerical solvers. SCALE: Simulation-Calibrated Amortized Learning for Energy Materials and MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback automate material and molecular design.
- Clinical and Engineering Applications: ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation and CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories provide real-time clinical support. Can Agents Design Better Chips with a Higher Level Abstraction? and PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking demonstrate AI’s role in complex engineering tasks.